Sphinx3增量索引合并到主索引

本文介绍了在CentOS 7系统下Sphinx-3.1.1的增量索引实现原理、模拟流程及定时任务配置。实现原理包括新建表记录ID、取新数据创建增量索引、合并索引等;模拟流程展示了数据库添加记录后生成和合并增量索引的过程;定时任务可按需设置运行频率。

摘要生成于 C知道 ,由 DeepSeek-R1 满血版支持, 前往体验 >

环境介绍

  1. 操作系统CentOS 7
  2. Sphinx-3.1.1

Sphinx的安装与配置,请看

https://blog.youkuaiyun.com/zhou16333/article/details/95131962

实现原理:

  1. 新建一张表,记录一下上一次已经创建好索引的最后一条记录的ID
  2. 当索引时,然后从数据库中取出所有ID大于上面那个sphinx中的那个ID的数据, 这些就是新的数据,然后创建一个小的索引文件
  3. 把上边我们创建的增量索引文件合并到主索引文件上去
  4. 把最后一条记录的ID更新到第一步创建的表中

值得注意的两点:

  1. 当合并索引的时候,只是把增量的索引合并进主索引中,增量索引本身并不会变化,也不会被删除;
  2. 当重建主索引的时候,增量索引就会被删除;

首先建立一个计数表,保存数据表的最新记录ID

CREATE TABLE `sph_counter` (
  `counter_id` int(11) NOT NULL COMMENT '标识不同的数据表',
  `max_doc_id` int(11) NOT NULL COMMENT '每个索引表的最大ID,会实时更新',
  PRIMARY KEY (`counter_id`)
) ENGINE=InnoDB DEFAULT CHARSET=utf8mb4;

配置索引文件
vim /application/sphinx/etc/test1.conf

#
# Minimal Sphinx configuration sample (clean, simple, functional)
#
# 主索引
source src1
{
	type			= mysql

	sql_host		= localhost
	sql_user		=root
	sql_pass		=root
	sql_db			= test
	sql_port		= 3306	# optional, default is 3306
	sql_query_pre		= SET NAMES utf8
	sql_query_pre		= REPLACE INTO sph_counter SELECT 1,MAX(id) FROM documents
	sql_query		= \
		SELECT id, group_id, UNIX_TIMESTAMP(date_added) AS date_added, title, content \
		FROM documents

	sql_attr_uint		= group_id
	sql_attr_timestamp	= date_added
}
index test1
{
	source			= src1
	path			= /sphinx/var/data/test1
}



# 增量索引
source src1_delta:src1
{
	sql_query_pre		= SET NAMES utf8
	sql_query		= \
		SELECT id, group_id, UNIX_TIMESTAMP(date_added) AS date_added, title, content FROM documents WHERE id > (SELECT IFNULL(max_doc_id,0) FROM sph_counter)
	sql_query_post		= REPLACE INTO sph_counter SELECT 1, MAX(id) FROM documents

	sql_attr_uint		= group_id
	sql_attr_timestamp	= date_added
}
index test1_delta:test1
{
        source                  = src1_delta
        path                    = /sphinx/var/data/test1_delta
}





index testrt
{
	type			= rt
	rt_mem_limit		= 128M

	path			= /sphinx/var/data/testrt

	rt_field		= title
	rt_field		= content
	rt_attr_uint		= gid
}


indexer
{
	mem_limit		= 128M
}


searchd
{
	listen			= 9312
	listen			= 9306:mysql41
	log			= /sphinx/var/log/searchd.log
	query_log		= /sphinx/var/log/query.log
	read_timeout		= 5
	max_children		= 30
	pid_file		= /sphinx/var/log/searchd.pid
	seamless_rotate		= 1
	preopen_indexes		= 1
	unlink_old		= 1
	workers			= threads # for RT to work
	binlog_path		= /sphinx/var/data
}

保存test1.conf配置文件后,先停止searchd进程,再启动
停止searchd
/application/sphinx/bin/searchd -c /application/sphinx/etc/test1.conf --stop
启动searchd
/application/sphinx/bin/searchd -c /application/sphinx/etc/test1.conf

关于索引的操作,下面三条命名只供观看,目前不要执行。

新建主索引
/application/sphinx/bin/indexer -c /application/sphinx/etc/test1.conf test1 --rotate
新建增量索引
/application/sphinx/bin/indexer -c /application/sphinx/etc/test1.conf test1_delta --rotate
合并索引
/application/sphinx/bin/indexer -c /application/sphinx/etc/test1.conf --merge test1 test1_delta --rotate

模拟流程

一、数据库已存在数据

mysql> SELECT * FROM documents;
+----+----------+-----------+---------------------+-----------------+---------------------------------------------------------------------------+
| id | group_id | group_id2 | date_added          | title           | content                                                                   |
+----+----------+-----------+---------------------+-----------------+---------------------------------------------------------------------------+
|  1 |        1 |         5 | 2019-07-08 10:56:34 | test one        | this is my test document number one. also checking search within phrases. |
|  2 |        1 |         6 | 2019-07-08 10:56:34 | test two        | this is my test document number two                                       |
|  3 |        2 |         7 | 2019-07-08 10:56:34 | another doc     | this is another group                                                     |
|  4 |        2 |         8 | 2019-07-08 10:56:34 | doc number four | this is to test groups                                                    |
+----+----------+-----------+---------------------+-----------------+---------------------------------------------------------------------------+
4 rows in set (0.00 sec)
mysql> SELECT * FROM sph_counter;
+------------+------------+
| counter_id | max_doc_id |
+------------+------------+
|          1 |          4 |
+------------+------------+
1 row in set (0.00 sec)

二、假设,系统经过某些操作之后,documents表中添加一条新记录。

+----+----------+-----------+---------------------+-----------------+---------------------------------------------------------------------------+
| id | group_id | group_id2 | date_added          | title           | content                                                                   |
+----+----------+-----------+---------------------+-----------------+---------------------------------------------------------------------------+
|  5 |        2 |         9 | 2019-07-09 16:52:52 | five test       | this is rows by good man                                                  |
+----+----------+-----------+---------------------+-----------------+---------------------------------------------------------------------------+

三、生成增量索引
/application/sphinx/bin/indexer -c /application/sphinx/etc/test1.conf test1_delta --rotate
此时

  1. 更新了增量索引 test1_delta
  2. sph_counter表数据为
mysql> SELECT * FROM sph_counter;
+------------+------------+
| counter_id | max_doc_id |
+------------+------------+
|          1 |          4 |
+------------+------------+

四、增量索引合并至主索引
/application/sphinx/bin/indexer -c /application/sphinx/etc/test1.conf --merge test1 test1_delta --rotate

定时任务

增量索引可以放在crontab里根据需要设置几分钟运行一次,然后执行索引合并

##每5分钟运行增量索引
*/5 * * * /application/sphinx/bin/indexer -c /application/sphinx/etc/csft.conf test1_delta --rotate > /dev/null 2>&1
##每10分钟执行一次增量索引合并
*/10 * * * /application/sphinx/bin/indexer -c /application/sphinx/etc/csft.conf --merge test1 test1_delta --rotate

##凌晨0点5分重新建立主索引
5 0 * * * /application/sphinx/bin/indexer -c /application/sphinx/etc/csft.conf --all --rotate > /dev/null 2>&1

参考文献

[1] [EB/OL]. https://www.cnblogs.com/mingaixin/p/5085708.html
[2] [EB/OL]. https://www.cnblogs.com/latma/p/6019834.html

评论 3
添加红包

请填写红包祝福语或标题

红包个数最小为10个

红包金额最低5元

当前余额3.43前往充值 >
需支付:10.00
成就一亿技术人!
领取后你会自动成为博主和红包主的粉丝 规则
hope_wisdom
发出的红包
实付
使用余额支付
点击重新获取
扫码支付
钱包余额 0

抵扣说明:

1.余额是钱包充值的虚拟货币,按照1:1的比例进行支付金额的抵扣。
2.余额无法直接购买下载,可以购买VIP、付费专栏及课程。

余额充值