Andy Pavlo joins ClickHouse to establish ClickHouse Labs
41–50 of 82 posts
Re: Andy Pavlo joins ClickHouse to establish ClickHouse Labs
#42Earlier quoted context omitted.
Many database companies have research labs. It allows you to explore ideas, unorthodox concepts, and parts of the design space that may never be reflected in an actual product. Or to figure out how to solve specific hard problems that you come across with neither a good solution nor proof of impossibility in literature. This is essential for a company that wants to stay on the frontier of database tech.
Which database companies have research labs? AFAIK amazon, snowflake, databricks, google and SAP don’t have a research lab dedicated to databases. They have some people that are paid to do research but there does not seem to be an IBM Almaden anywhere in the world.
Re: Andy Pavlo joins ClickHouse to establish ClickHouse Labs
#43Earlier quoted context omitted.
I'd ascribe "deep tech" to anything that you can reasonably get a PhD in and have it not be unusual. There are dozens of academic conferences on DBs pushing the frontier forward.
Don't feed the troll. The dude doesn't know anything about databases.
However, the commenter I was responding to apparently considers just about _any_ research in any significant field to be deep tech so w/e I give up and I'll go update my CV to include the phrase "deep tech research"
Re: Andy Pavlo joins ClickHouse to establish ClickHouse Labs
#44I'm very curious about the convergence of the best in class fast OLAP products (StarRocks, ClickHouse) with Trino. It sounds like everybody is going for decoupled compute/storage, using S3 or similar as the storage layer, and thus forgoing colocated joins (ok I know ClickHouse joins suck)... So what does this mean for ingestion (and indexing)? Iceberg V3? Paimon? Bespoke ingestion through the DB engine to do the inde…
It's also interesting how Clickhouse / Starrocks can now also act as a query planner and executor on top of non-native formats (ex. Iceberg). I assume the native formats will always be faster / more optimized but the need for Trino as a separate executor while running either of these databases seems to be close to gone.
Native format is faster (especially for colocated joins), but it's way more expensive if you have to run a bunch of separate storage nodes vs just using S3, especially your query volume isn't that high.
I liken it to the BigQuery cost model, where storage is effectively free.
Re: Andy Pavlo joins ClickHouse to establish ClickHouse Labs
#45Earlier quoted context omitted.
Deep tech fields have a more complex relationship with academia than other industries. Engineers often need license to test ideas that may not see adoption for years, or require tens of millions to build-out. A "research" arm is exactly this license, although it comes at the cost of potentially killing innovation in the rest of the company.
C'mon, databases are great but they are not in any way a "deep tech" field. Databases exist today in a zillion production forms as commercial products and free software. Deep tech includes things like nuclear fusion, solid state batteries, quantum computers. I know everyone wants to feel cool, but just because your new javascript framework will be in beta for the next ten years doesn't make it "deep tech".
And I bet DB engines can go far than normal OS.
You don't know how much is still waiting for somebody to try, and how much is not applied. And how many of that DBs are not even doing, because all are constrained be being "apps" with so poor interface (sql).
Fun fact: Not exist a viable true relational DBs implemented, neither exist one with a viable programming language AND apis that is for developers.
ZERO.
Re: Andy Pavlo joins ClickHouse to establish ClickHouse Labs
#46Earlier quoted context omitted.
The databases we use today in production have severe limitations and are not even close to what is theoretically possible. Many traditional parts of a database (indexing, caching, scheduling, et al) are AI-complete algorithm problems. Entire sub-classes of database (e.g. graph or spatial) famously have persistently poor scalability and performance because of open questions in the foundational computer science. Just t…
Could you explain what practical research there is to be done? The heavy theory I know does not seem to be very useful in practice. Optimal join algorithms, Yannakakis adjacent algorithms, tree decomposition of queries all seem to be worse than well implemented naive algorithms. But maybe the implementations of the new algorithms just are not good? I really don’t know.
Ideally a table should be index-organized across all relevant columns. No public system works anything like this. We don't have single indexing structures that work for a collection of arbitrary types each with possibly unpredictable distributions, never mind ones that mix temporal, geometric, and other difficult types. The AI-complete nature of indexing becomes evident when you dig into this. Downstream from this is an implication of extremely granular and adaptive storage management that current storage engines aren't designed for.
Tractable cache replacement algorithms are broken for many workloads and data models. These algorithms need to be very fast for search, update, and eviction selection but they are also AI-complete; improvements to generality have impractically high computational cost. Storage growth is decoupled from RAM availability thanks to disaggregation, aggravating the problem even for workloads that worked well under tractable cache replacement. In theory we know that cache admission (read: fancy latency-hiding schedules) is more robust and scales better but is so difficult to implement in non-trivial real systems that I don't think anyone has figured out how to reduce that concept to practice yet.
At exabyte scales, conventional database internals have embedded assumptions that no longer hold true. For example, you cannot guarantee even "small" internal control structures are resident in memory. What used to be fairly boring internals bits in databases suddenly have to be redesigned from first principles. This is more applied than theoretical but it suggests a major change in the way we do internal architecture.
Traditionally we've treated spatial and temporal locality as architecturally separate concerns. This is extremely convenient from a building real systems standpoint. Optimizing either one in isolation is adversarial to the efficiency of the other, which becomes increasingly visible as you scale up. Converging these concerns into a single "thing" almost certainly has solving the above problems as a prerequisite. If you squint, you can kind of see this as the last step before databases become literal AI.
All of these have really broad scope. If we could solve even half of these open research problems the resulting database engines would be unrecognizable. There are ton of other narrower interesting research problems around data layouts, compression, join parallelism, etc that still have potential for substantial improvement.
It is a great time to be doing database research, we've barely scratched the surface.
Re: Andy Pavlo joins ClickHouse to establish ClickHouse Labs
#47Earlier quoted context omitted.
The databases we use today in production have severe limitations and are not even close to what is theoretically possible. Many traditional parts of a database (indexing, caching, scheduling, et al) are AI-complete algorithm problems. Entire sub-classes of database (e.g. graph or spatial) famously have persistently poor scalability and performance because of open questions in the foundational computer science. Just t…
Could you explain what practical research there is to be done? The heavy theory I know does not seem to be very useful in practice. Optimal join algorithms, Yannakakis adjacent algorithms, tree decomposition of queries all seem to be worse than well implemented naive algorithms. But maybe the implementations of the new algorithms just are not good? I really don’t know.
Are there re-usable query primitives for extremely large scale multi-modal data? how do you scale such queries or make them efficient?
The list goes on.
Re: Andy Pavlo joins ClickHouse to establish ClickHouse Labs
#48Earlier quoted context omitted.
C'mon, databases are great but they are not in any way a "deep tech" field. Databases exist today in a zillion production forms as commercial products and free software. Deep tech includes things like nuclear fusion, solid state batteries, quantum computers. I know everyone wants to feel cool, but just because your new javascript framework will be in beta for the next ten years doesn't make it "deep tech".
Maybe only an OS can be close to how MUCH deep you can go with a DB engine. And I bet DB engines can go far than normal OS. You don't know how much is still waiting for somebody to try, and how much is not applied. And how many of that DBs are not even doing, because all are constrained be being "apps" with so poor interface (sql). Fun fact: Not exist a viable true relational DBs implemented, neither exist one with a…
but that's not what this term has historically meant https://en.wikipedia.org/wiki/Deep_tech
> Deep tech innovations are often radical and may create new markets or disrupt existing ones. Deep tech companies often address big societal and environmental challenges and have potential to impact everyday life. Silicon chips are an example of innovation that enabled calculation at previously unimaginable speed and scale.
Database research is good, important, critical, even! But it's not creating something new that has never existed before. It's not inventing the transistor or the integrated circuit.
Re: Andy Pavlo joins ClickHouse to establish ClickHouse Labs
#49Always enjoyed his lecture series from CMU, hopefully those continue in a sponsored format from Clickhouse.
They will continue. New seminar series starts next month (announcement coming this week).
Re: Andy Pavlo joins ClickHouse to establish ClickHouse Labs
#50Earlier quoted context omitted.
What do you mean? There’s plenty of high impact research on databases happening in academia. (Including, for example, Andy’s group at CMU) Also keep in mind that the part you quoted is partially marketing copy.
I take this announcement to mean he is leaving. So he's a perfect example
It's a good thing when industry is a competitively attractive environment for research, which is how I read this.