Nutch2.1+mysql+solr3.6.1+中文网站抓取

最新推荐文章于 2021-01-28 09:35:16 发布

转载最新推荐文章于 2021-01-28 09:35:16 发布 · 1.1k 阅读

nutch 专栏收录该内容

6 篇文章

订阅专栏

本文介绍如何配置Nutch爬虫抓取网页，并使用Solr进行内容索引。主要内容包括Nutch与MySQL数据库的集成、Solr的安装与配置、爬虫抓取流程及结果验证。

摘要生成于 C知道，由 DeepSeek-R1 满血版支持，前往体验 >

1、mysql 数据库配置

linux mysql安装步骤省略。

创建数据库与表

[sql] view plain copy print ?

CREATE DATABASE nutch DEFAULT CHARACTER SET utf8 DEFAULT COLLATE utf8_general_ci;
CREATE TABLE `webpage` (
`id` varchar(767) CHARACTER SET latin1 NOT NULL,
`headers` blob,
`text` mediumtext DEFAULT NULL,
`status` int(11) DEFAULT NULL,
`markers` blob,
`parseStatus` blob,
`modifiedTime` bigint(20) DEFAULT NULL,
`score` float DEFAULT NULL,
`typ` varchar(32) CHARACTER SET latin1 DEFAULT NULL,
`baseUrl` varchar(512) CHARACTER SET latin1 DEFAULT NULL,
`content` mediumblob,
`title` varchar(2048) DEFAULT NULL,
`reprUrl` varchar(512) CHARACTER SET latin1 DEFAULT NULL,
`fetchInterval` int(11) DEFAULT NULL,
`prevFetchTime` bigint(20) DEFAULT NULL,
`inlinks` mediumblob,
`prevSignature` blob,
`outlinks` mediumblob,
`fetchTime` bigint(20) DEFAULT NULL,
`retriesSinceFetch` int(11) DEFAULT NULL,
`protocolStatus` blob,
`signature` blob,
`metadata` blob,
PRIMARY KEY (`id`)
) ENGINE=InnoDB DEFAULT CHARSET=utf8;

2、安装nutch2.1
A、 nutch下载地址：http://apache.etoak.com/nutch/2.1/apache-nutch-2.1-src.zip

下载完成后家压缩，

B、以下将nutch的根目录定位${APACHE_NUTCH_HOME}.

C、配置nutch对mysql的支持，修改${APACHE_NUTCH_HOME}/ivy/ivy.xml文件

将这行的注释取消<dependency org=”mysql” name=”mysql-connector-java” rev=”5.1.18″ conf=”*->default”/>

修改${APACHE_NUTCH_HOME}/conf/gora.properties文件，

注释默认存储配置

[html] view plain copy print ?

###############################
# Default SqlStore properties #
###############################
#gora.sqlstore.jdbc.driver=org.hsqldb.jdbc.JDBCDriver
#gora.sqlstore.jdbc.url=jdbc:hsqldb:hsql://localhost/nutchtest
#gora.sqlstore.jdbc.user=sa
#gora.sqlstore.jdbc.password=
取消以下代码注释，
###############################
# MySQL properties
################################
gora.sqlstore.jdbc.driver=com.mysql.jdbc.Driver
gora.sqlstore.jdbc.url=jdbc:mysql://localhost:3306/nutch?createDatabaseIfNotExist=true
gora.sqlstore.jdbc.user=xxxxx（mysql用户名）
gora.sqlstore.jdbc.password=xxxxx（mysql密码）

D、修改${APACHE_NUTCH_HOME}/conf/nutch-site.xml 加入如下代码：

[html] view plain copy print ?

<property>

<name>http.agent.name</name>

<value>Your Nutch Spider</value>

</property>

<property>

<name>http.accept.language</name>

<value>ja-jp, en-us,en-gb,en;q=0.7,*;q=0.3</value>

<description>Value of the “Accept-Language” request header field.

This allows selecting non-English language as default one to retrieve.

It is a useful setting for search engines build for certain national group.

</description>

</property>

<property>

<name>parser.character.encoding.default</name>

<value>utf-8</value>

<description>The character encoding to fall back to when no other information

is available</description>

</property>

<property>

<name>storage.data.store.class</name>

<value>org.apache.gora.sql.store.SqlStore</value>

<description>The Gora DataStore class for storing and retrieving data.

Currently the following stores are available: ….

</description>

</property>