初学正则表达式 -- Java

最新推荐文章于 2025-09-12 16:46:38 发布

xiaosi23

最新推荐文章于 2025-09-12 16:46:38 发布

阅读量563

点赞数

CC 4.0 BY-SA版权

文章标签：正则表达式 html microsoft attributes class string

本文链接：https://blog.youkuaiyun.com/xiaosi23/article/details/1726471

本文分享了作者在使用正则表达式过程中的一些实用技巧，包括如何匹配特殊字符、使用特殊符号以及处理文本中特定模式的方法。此外还提供了一些去除HTML标签的正则表达式示例。

基本介绍请参照 http://www.blogjava.net/lbx19822004/archive/2007/05/28/120423.html
写的比较详细了
也可以参考http://msdn.microsoft.com/library/chs/default.asp?url=/library/CHS/jscript7/html/jsjsgrpregexpsyntax.asp

这里就我个人试验的结果进行一些补充说明:
1) "." 匹配除了换行符以外的所有字符, 包括空格, 但是不包括换行符. 这点一定要注意
2) 对于新手, 正则表达式的使用经常会有一些很灵异的现象发生, 例如"." 和 "[.]" 就不是一样的效果, 具体原因,俺还不清楚
3) 如果要匹配包括换行符在内的所有字符, 可以考虑使用"(.|//s)" ,但是"[.//s]"和"[.|//s]"不行,也不知道为什么.
4) 注意"^"的用法: 它有两种用法: 1."[^]",当他与"["搭配使用时,代表除此之外, 例如: [^abc] 代表除了a,b,c以外的其他所有字符,请注意,这里不是字符串, 请不要误解为: 除了字符串"abc"以外的其他字符 . 2.代表一行的开始
5) 如果想要去掉文档中的连续多个空行, 可以考虑,将"//s+/r/n" 替换为"/r/n"

暂时就用到这么多, 以后再总结哈

补充2:
对于以上的3) 可以使用"[//s//S]"
新发现: http://tim.mackey.ie/CleanWordHTMLUsingRegularExpressions.aspx
两个很有用的正则表达式

/// <summary>
/// Removes all FONT and SPAN tags, and all Class and Style attributes.
/// Designed to get rid of non-standard Microsoft Word HTML tags.
/// </summary>
private string CleanHtml(string html)
{ 
    // start by completely removing all unwanted tags 
    html = Regex.Replace(html, @"<[/]?(font|span|xml|del|ins|[ovwxp]:/w+)[^>]*?>", "", RegexOptions.IgnoreCase); 
    // then run another pass over the html (twice), removing unwanted attributes 
    html = Regex.Replace(html, @"<([^>]*)(?:class|lang|style|size|face|[ovwxp]:/w+)=(?:'[^']*'|""[^""]*""|[^/s>]+)([^>]*)>","<$1$2>", RegexOptions.IgnoreCase); 
    html = Regex.Replace(html, @"<([^>]*)(?:class|lang|style|size|face|[ovwxp]:/w+)=(?:'[^']*'|""[^""]*""|[^/s>]+)([^>]*)>","<$1$2>", RegexOptions.IgnoreCase); 
    return html;
}