Groonga開源搜尋引擎——列儲存做聚合，沒有內建分散式，分片和副本是隨mysql或者postgreSQL作為儲存引擎由MySQL自身來做分片和副本的...

weixin_34119545發表於2017-05-03

1. Characteristics of Groonga

ppt：http://mroonga.org/publication/presentation/groonga-mysqluc2011.pdf

1.1. Groonga overview

Groonga is a fast and accurate full text search engine based on inverted index. One of the characteristics of Groonga is that a newly registered document instantly appears in search results. Also, Groonga allows updates without read locks. These characteristics result in superior performance on real-time applications.

Groonga is also a column-oriented database management system (DBMS). Compared with well-known row-oriented systems, such as MySQL and PostgreSQL, column-oriented systems are more suited for aggregate queries. Due to this advantage, Groonga can cover weakness of row-oriented systems.

The basic functions of Groonga are provided in a C library. Also, libraries for using Groonga in other languages, such as Ruby, are provided by related projects. In addition, groonga-based storage engines are provided for MySQL and PostgreSQL. These libraries and storage engines allow any application to use Groonga. See usage examples.

1.2. Full text search and Instant update

In widely used DBMSs, updates are immediately processed, for example, a newly registered record appears in the result of the next query. In contrast, some full text search engines do not support instant updates, because it is difficult to dynamically update inverted indexes, the underlying data structure.

Groonga also uses inverted indexes but supports instant updates. In addition, Groonga allows you to search documents even when updating the document collection. Due to these superior characteristics, Groonga is very flexible as a full text search engine. Also, Groonga always shows good performance because it divides a large task, inverted index merging, into smaller tasks.

1.3. Column store and aggregate query

People can collect more than enough data in the Internet era. However, it is difficult to extract informative knowledge from a large database, and such a task requires a many-sided analysis through trial and error. For example, search refinement by date, time and location may reveal hidden patterns. Aggregate queries are useful to perform this kind of tasks.

An aggregate query groups search results by specified column values and then counts the number of records in each group. For example, an aggregate query in which a location column is specified counts the number of records per location. Making a graph from the result of an aggregate query against a date column is an easy way to visualize changes over time. Also, a combination of refinement by location and an aggregate query against a date column allows visualization of changes over time in specific location. Thus refinement and aggregation are important to perform data mining.

A column-oriented architecture allows Groonga to efficiently process aggregate queries because a column-oriented database, which stores records by column, allows an aggregate query to access only a specified column. On the other hand, an aggregate query on a row-oriented database, which stores records by row, has to access neighbor columns, even though those columns are not required.

1.4. Inverted index and tokenizer

An inverted index is a traditional data structure used for large-scale full text search. A search engine based on inverted index extracts index terms from a document when it is added. Then in retrieval, a query is divided into index terms to find documents containing those index terms. In this way, index terms play an important role in full text search and thus the way of extracting index terms is a key to a better search engine.

A tokenizer is a module to extract index terms. A Japanese full text search engine commonly uses a word-based tokenizer (hereafter referred to as a word tokenizer) and/or a character-based n-gram tokenizer (hereafter referred to as an n-gram tokenizer). A word tokenizer-based search engine is superior in time, space and precision, which is the fraction of relevant documents in a search result. On the other hand, an n-gram tokenizer-based search engine is superior in recall, which is the fraction of retrieved documents in the perfect search result. The best choice depends on the application in practice.

Groonga supports both word and n-gram tokenizers. The simplest built-in tokenizer uses spaces as word delimiters. Built-in n-gram tokenizers (n = 1, 2, 3) are also available by default. In addition, a yet another built-in word tokenizer is available if MeCab, a part-of-speech and morphological analyzer, is embedded. Note that a tokenizer is pluggable and you can develop your own tokenizer, such as a tokenizer based on another part-of-speech tagger or a named-entity recognizer.

1.5. Sharable storage and read lock-free

Multi-core processors are mainstream today and the number of cores per processor is increasing. In order to exploit multiple cores, executing multiple queries in parallel or dividing a query into sub-queries for parallel processing is becoming more important.

A database of Groonga can be shared with multiple threads/processes. Also, multiple threads/processes can execute read queries in parallel even when another thread/process is executing an update query because Groonga uses read lock-free data structures. This feature is suited to a real-time application that needs to update a database while executing read queries. In addition, Groonga allows you to build flexible systems. For example, a database can receive read queries through the built-in HTTP server of Groonga while accepting update queries through MySQL.

1.6. Geo-location (latitude and longitude) search

Location services are getting more convenient because of mobile devices with GPS. For example, if you are going to have lunch or dinner at a nearby restaurant, a local search service for restaurants may be very useful, and for such services, fast geo-location search is becoming more important.

Groonga provides inverted index-based fast geo-location search, which supports a query to find points in a rectangle or circle. Groonga gives high priority to points near the center of an area. Also, Groonga supports distance measurement and you can sort points by distance from any point.

1.7. Groonga library

The basic functions of Groonga are provided in a C library and any application can use Groonga as a full text search engine or a column-oriented database. Also, libraries for languages other than C/C++, such as Ruby, are provided in related projects. See related projects for details.

1.8. Groonga server

Groonga provides a built-in server command which supports HTTP, the memcached binary protocol and the Groonga Query Transfer Protocol (GQTP). Also, a Groonga server supports query caching, which significantly reduces response time for repeated read queries. Using this command, Groonga is available even on a server that does not allow you to install new libraries.

1.9. Mroonga storage engine

Groonga works not only as an independent column-oriented DBMS but also as storage engines of well-known DBMSs. For example, Mroonga is a MySQL pluggable storage engine using Groonga. By using Mroonga, you can use Groonga for column-oriented storage and full text search. A combination of a built-in storage engine, MyISAM or InnoDB, and a Groonga-based full text search engine is also available. All the combinations have good and bad points and the best one depends on the application. See related projects for details.

轉自：http://groonga.org/docs/characteristic.html

待分析！

MySQL 儲存引擎
2020-08-29
MySql儲存引擎
MySQL儲存引擎
2024-08-23
MySql儲存引擎
MySQL系列-儲存引擎
2018-12-03
MySql儲存引擎
MySQL Archive儲存引擎
2016-08-30
MySqlHive儲存引擎
MySql 官方儲存引擎
2017-03-25
MySql儲存引擎
MySQL MEMORY儲存引擎
2014-12-22
MySql儲存引擎
MySQL InnoDB儲存引擎
2024-05-25
MySql儲存引擎
Mysql技術內幕InnoDB儲存引擎讀書筆記--《一》Mysql體系結構和儲存引擎
2017-06-30
MySql儲存引擎筆記
理解mysql的儲存引擎
2021-09-09
MySql儲存引擎
MySql體系結構和儲存引擎
2017-03-10
MySql儲存引擎
MySQL儲存引擎：MyISAM和InnoDB的區別
2020-12-09
MySql儲存引擎
【Mysql 學習】Mysql 儲存引擎
2011-01-05
MySql儲存引擎
MySQL入門--儲存引擎
2019-06-28
MySql儲存引擎
MySQL之四儲存引擎
2021-03-02
MySql儲存引擎
(5)mysql 常用儲存引擎
2017-01-18
MySql儲存引擎
MySQL-05.儲存引擎
2024-04-19
MySql儲存引擎
[Mysql技術內幕]Innodb儲存引擎
2021-05-19
MySql儲存引擎
MySQL技術內幕:InnoDB儲存引擎
2011-06-08
MySql儲存引擎
openGauss儲存技術（二）——列儲存引擎和記憶體引擎
2022-11-09
儲存引擎記憶體
Mysql技術內幕InnoDB儲存引擎讀書筆記--《二》InnoDB儲存引擎
2017-06-30
MySql儲存引擎筆記
MySQL InnoDB 儲存引擎探祕
2019-02-21
MySql儲存引擎
2_mysql（索引、儲存引擎）
2020-11-15
MySql索引儲存引擎
MySQL federated儲存引擎測試
2023-01-13
MySql儲存引擎
MySql 擴充套件儲存引擎
2017-03-25
MySql套件儲存引擎
MySQL 5.5儲存引擎介紹
2016-04-13
MySql儲存引擎
【Mysql 學習】memory儲存引擎
2011-01-05
MySql儲存引擎
MySQL 資料庫儲存引擎
2015-11-11
MySql資料庫儲存引擎
聊一聊MySQL的儲存引擎
2022-01-24
MySql儲存引擎
如何選擇mysql的儲存引擎
2018-10-11
MySql儲存引擎
MyISAM 儲存引擎,Innodb 儲存引擎
2015-02-28
儲存引擎
MySQL2：四種MySQL儲存引擎
2015-11-07
MySql儲存引擎
Mysql 的儲存過程和儲存函式
2013-08-02
MySql儲存過程儲存函式
InnoDB 作為預設儲存引擎（從mysql-5.5.5開始)薦
2011-01-12
儲存引擎MySql
mysql和orcale的儲存過程和儲存函式
2018-06-04
MySql儲存過程儲存函式
【Mysql技術內幕筆記--1】--Mysql體系結構和儲存引擎
2021-09-09
MySql筆記儲存引擎
《MySQL技術內幕:InnoDB儲存引擎》連載
2011-04-06
MySql儲存引擎
小談mysql儲存引擎優化
2018-07-16
MySql儲存引擎優化
MySQL儲存引擎入門介紹
2020-09-28
MySql儲存引擎