I did something similar for a NoSQL database [1]. The biggest surprise was how much the query performance can change for an index when the data distribution changes slightly. For example using a real distribution for an 'age' field instead of just using a random number like in the test data. [1] https://rxdb.info/query-optimizer.html
The SQLite Index Suggester
11–20 of 26 posts
Re: The SQLite Index Suggester
#12I find it interesting, as recently I was reading about how complex it would/could be to create an index suggester. https://www.depesz.com/2021/10/22/why-is-it-hard-to-automati...
This suggester isn’t very good. It takes a single query and suggests indexes for it. A good one would take a mix of queries and suggest a set of indexes, also considering the impact on write speed of additional indexes (table updates often need to update indexes, too)
For the example in this article, if the table is large and the average number of rows with a given ‘a’ value is close to 1 or if most queries are for ‘a’ values that aren’t in the database, it may even be better to do
CREATE INDEX x1a ON x1(a);
That gives you a smaller index, decreasing disk usage.Re: The SQLite Index Suggester
#13This sounds really cool! I've sometimes wondered why server-based RDBMSs don't offer something like this. Is it too hard to implement? Or did people just not think of it? Or do they have something like this and I just never learned about it?
Re: The SQLite Index Suggester
#14This sounds really cool! I've sometimes wondered why server-based RDBMSs don't offer something like this. Is it too hard to implement? Or did people just not think of it? Or do they have something like this and I just never learned about it?
Re: The SQLite Index Suggester
#15This sounds really cool! I've sometimes wondered why server-based RDBMSs don't offer something like this. Is it too hard to implement? Or did people just not think of it? Or do they have something like this and I just never learned about it?
Microsoft SQL Server definitely has suggestions for missing indexes. The quality of the suggestions are debatable though
I don't know if it's still around, but in mid-2000s it was light years ahead of any other database.
Re: The SQLite Index Suggester
#16I find it interesting, as recently I was reading about how complex it would/could be to create an index suggester. https://www.depesz.com/2021/10/22/why-is-it-hard-to-automati...
It’s like writing a compiler or interpreter: writing one is easy; writing a good one extremely hard. This suggester isn’t very good. It takes a single query and suggests indexes for it. A good one would take a mix of queries and suggest a set of indexes, also considering the impact on write speed of additional indexes (table updates often need to update indexes, too) For the example in this article, if the table is l…
The underlying API can analyse multiple queries - looks like they've only coded up the test `.expert` command for one.
From [1], "The sqlite3expert object is configured with one or more SQL statements by making one or more calls to sqlite3_expert_sql(). Each call may specify a single SQL statement, or multiple statements separated by semi-colons." then "sqlite3_expert_analyze() is called to run the analysis."
Re: The SQLite Index Suggester
#17I find it interesting, as recently I was reading about how complex it would/could be to create an index suggester. https://www.depesz.com/2021/10/22/why-is-it-hard-to-automati...
It’s like writing a compiler or interpreter: writing one is easy; writing a good one extremely hard. This suggester isn’t very good. It takes a single query and suggests indexes for it. A good one would take a mix of queries and suggest a set of indexes, also considering the impact on write speed of additional indexes (table updates often need to update indexes, too) For the example in this article, if the table is l…
Since we wrote our initial index suggestion tool for Postgres, we actually went back to the drawing board, examined the concerns brought up, and developed a new per-table Index Advisor for Postgres that we recently released [1].
The gist of it: Instead of looking at the "perfect" index for each query, its important to test out different "good enough" indexes that cover multiple queries. Additionally, as you note, the write overhead of indexes needs to be considered (both from a table writes / second approach, as well as disk space used at a given moment in time).
I think this is a fascinating field and there is lots more work to be done. I've also found the 2020 paper "Experimental Evaluation of Index Selection Algorithms" [2] pretty useful, that compares a few different approaches.
[1] https://pganalyze.com/blog/automatic-indexing-system-postgre...
Re: The SQLite Index Suggester
#18I did something similar for a NoSQL database [1]. The biggest surprise was how much the query performance can change for an index when the data distribution changes slightly. For example using a real distribution for an 'age' field instead of just using a random number like in the test data. [1] https://rxdb.info/query-optimizer.html
Re: The SQLite Index Suggester
#19I did something similar for a NoSQL database [1]. The biggest surprise was how much the query performance can change for an index when the data distribution changes slightly. For example using a real distribution for an 'age' field instead of just using a random number like in the test data. [1] https://rxdb.info/query-optimizer.html
It seems better for birthdate to be stored in the database and age just to be calculated when needed?
Re: The SQLite Index Suggester
#20Earlier quoted context omitted.
Microsoft SQL Server definitely has suggestions for missing indexes. The quality of the suggestions are debatable though
Microsoft SQL Server query analyzer was was essential in identifying missing indexes. Wherever you saw "full table scan" on a table, you knew it was missing an index. I don't know if it's still around, but in mid-2000s it was light years ahead of any other database.