环境介绍
- 操作系统CentOS 7
- Sphinx-3.1.1
Sphinx的安装与配置,请看
https://blog.youkuaiyun.com/zhou16333/article/details/95131962
实现原理:
- 新建一张表,记录一下上一次已经创建好索引的最后一条记录的ID
- 当索引时,然后从数据库中取出所有ID大于上面那个sphinx中的那个ID的数据, 这些就是新的数据,然后创建一个小的索引文件
- 把上边我们创建的增量索引文件合并到主索引文件上去
- 把最后一条记录的ID更新到第一步创建的表中
值得注意的两点:
- 当合并索引的时候,只是把增量的索引合并进主索引中,增量索引本身并不会变化,也不会被删除;
- 当重建主索引的时候,增量索引就会被删除;
首先建立一个计数表,保存数据表的最新记录ID
CREATE TABLE `sph_counter` (
`counter_id` int(11) NOT NULL COMMENT '标识不同的数据表',
`max_doc_id` int(11) NOT NULL COMMENT '每个索引表的最大ID,会实时更新',
PRIMARY KEY (`counter_id`)
) ENGINE=InnoDB DEFAULT CHARSET=utf8mb4;
配置索引文件
vim /application/sphinx/etc/test1.conf
#
# Minimal Sphinx configuration sample (clean, simple, functional)
#
# 主索引
source src1
{
type = mysql
sql_host = localhost
sql_user =root
sql_pass =root
sql_db = test
sql_port = 3306 # optional, default is 3306
sql_query_pre = SET NAMES utf8
sql_query_pre = REPLACE INTO sph_counter SELECT 1,MAX(id) FROM documents
sql_query = \
SELECT id, group_id, UNIX_TIMESTAMP(date_added) AS date_added, title, content \
FROM documents
sql_attr_uint = group_id
sql_attr_timestamp = date_added
}
index test1
{
source = src1
path = /sphinx/var/data/test1
}
# 增量索引
source src1_delta:src1
{
sql_query_pre = SET NAMES utf8
sql_query = \
SELECT id, group_id, UNIX_TIMESTAMP(date_added) AS date_added, title, content FROM documents WHERE id > (SELECT IFNULL(max_doc_id,0) FROM sph_counter)
sql_query_post = REPLACE INTO sph_counter SELECT 1, MAX(id) FROM documents
sql_attr_uint = group_id
sql_attr_timestamp = date_added
}
index test1_delta:test1
{
source = src1_delta
path = /sphinx/var/data/test1_delta
}
index testrt
{
type = rt
rt_mem_limit = 128M
path = /sphinx/var/data/testrt
rt_field = title
rt_field = content
rt_attr_uint = gid
}
indexer
{
mem_limit = 128M
}
searchd
{
listen = 9312
listen = 9306:mysql41
log = /sphinx/var/log/searchd.log
query_log = /sphinx/var/log/query.log
read_timeout = 5
max_children = 30
pid_file = /sphinx/var/log/searchd.pid
seamless_rotate = 1
preopen_indexes = 1
unlink_old = 1
workers = threads # for RT to work
binlog_path = /sphinx/var/data
}
保存test1.conf配置文件后,先停止searchd进程,再启动
停止searchd
/application/sphinx/bin/searchd -c /application/sphinx/etc/test1.conf --stop
启动searchd
/application/sphinx/bin/searchd -c /application/sphinx/etc/test1.conf
关于索引的操作,下面三条命名只供观看,目前不要执行。
新建主索引
/application/sphinx/bin/indexer -c /application/sphinx/etc/test1.conf test1 --rotate
新建增量索引
/application/sphinx/bin/indexer -c /application/sphinx/etc/test1.conf test1_delta --rotate
合并索引
/application/sphinx/bin/indexer -c /application/sphinx/etc/test1.conf --merge test1 test1_delta --rotate
模拟流程
一、数据库已存在数据
mysql> SELECT * FROM documents;
+----+----------+-----------+---------------------+-----------------+---------------------------------------------------------------------------+
| id | group_id | group_id2 | date_added | title | content |
+----+----------+-----------+---------------------+-----------------+---------------------------------------------------------------------------+
| 1 | 1 | 5 | 2019-07-08 10:56:34 | test one | this is my test document number one. also checking search within phrases. |
| 2 | 1 | 6 | 2019-07-08 10:56:34 | test two | this is my test document number two |
| 3 | 2 | 7 | 2019-07-08 10:56:34 | another doc | this is another group |
| 4 | 2 | 8 | 2019-07-08 10:56:34 | doc number four | this is to test groups |
+----+----------+-----------+---------------------+-----------------+---------------------------------------------------------------------------+
4 rows in set (0.00 sec)
mysql> SELECT * FROM sph_counter;
+------------+------------+
| counter_id | max_doc_id |
+------------+------------+
| 1 | 4 |
+------------+------------+
1 row in set (0.00 sec)
二、假设,系统经过某些操作之后,documents表中添加一条新记录。
+----+----------+-----------+---------------------+-----------------+---------------------------------------------------------------------------+
| id | group_id | group_id2 | date_added | title | content |
+----+----------+-----------+---------------------+-----------------+---------------------------------------------------------------------------+
| 5 | 2 | 9 | 2019-07-09 16:52:52 | five test | this is rows by good man |
+----+----------+-----------+---------------------+-----------------+---------------------------------------------------------------------------+
三、生成增量索引
/application/sphinx/bin/indexer -c /application/sphinx/etc/test1.conf test1_delta --rotate
此时
- 更新了增量索引 test1_delta
- sph_counter表数据为
mysql> SELECT * FROM sph_counter;
+------------+------------+
| counter_id | max_doc_id |
+------------+------------+
| 1 | 4 |
+------------+------------+
四、增量索引合并至主索引
/application/sphinx/bin/indexer -c /application/sphinx/etc/test1.conf --merge test1 test1_delta --rotate
定时任务
增量索引可以放在crontab里根据需要设置几分钟运行一次,然后执行索引合并
##每5分钟运行增量索引
*/5 * * * /application/sphinx/bin/indexer -c /application/sphinx/etc/csft.conf test1_delta --rotate > /dev/null 2>&1
##每10分钟执行一次增量索引合并
*/10 * * * /application/sphinx/bin/indexer -c /application/sphinx/etc/csft.conf --merge test1 test1_delta --rotate
##凌晨0点5分重新建立主索引
5 0 * * * /application/sphinx/bin/indexer -c /application/sphinx/etc/csft.conf --all --rotate > /dev/null 2>&1
参考文献
[1] [EB/OL]. https://www.cnblogs.com/mingaixin/p/5085708.html
[2] [EB/OL]. https://www.cnblogs.com/latma/p/6019834.html