IKAnalyzer中文分词

  1. Project structure

  2. IKAnalyzer.cfg.xml

    The configuration file of IKAnalyzer, a sentence separator supported Chinese. It must in the root of src.

    <?xml version="1.0" encoding="UTF-8"?>
    <!DOCTYPE properties SYSTEM "http://java.sun.com/dtd/properties.dtd">  
    <properties>  
    <comment>IK Analyzer extra configuration</comment>
    <!-- configure your own dic here -->
    <entry key="ext_dict">cn/com/tragicEnding/prov/util/ext.dic;</entry> 
    <!-- configure your own stop dic here -->
    <entry key="ext_stopwords">cn/com/tragicEnding/prov/util/stopword.dic</entry> 
    </properties>

  3. ext.dic / stopword.dic

    View the spec of dic.
        

  4. Keywords.java

    Function to separate sentence.

    public static List<String> splitToKeywords(String word)
    {
    	List<String> keywords = new ArrayList<String>();
    	try
    	{
    		Analyzer anal = new IKAnalyzer(true);
    		StringReader reader = new StringReader(word);
    		TokenStream ts = anal.tokenStream("", reader);
    		CharTermAttribute term = ts.getAttribute(CharTermAttribute.class);
    		while(ts.incrementToken())
    		{
    			keywords.add(term.toString());
    		}
    		reader.close();
    		anal.close();
    
    	}
    	catch(Exception e)
    	{
    		e.printStackTrace();
    	}
    
    	return keywords;
    }



评论
添加红包

请填写红包祝福语或标题

红包个数最小为10个

红包金额最低5元

当前余额3.43前往充值 >
需支付:10.00
成就一亿技术人!
领取后你会自动成为博主和红包主的粉丝 规则
hope_wisdom
发出的红包
实付
使用余额支付
点击重新获取
扫码支付
钱包余额 0

抵扣说明:

1.余额是钱包充值的虚拟货币,按照1:1的比例进行支付金额的抵扣。
2.余额无法直接购买下载,可以购买VIP、付费专栏及课程。

余额充值