Adrien Grand

@jpountz.bsky.social

#Lucene developer

It's interesting how the Elasticsearch and Datadog (www.datadoghq.com/blog/enginee...) approaches to wildcard search differ. Both use n-gram indexes, but with different strategies to contain storage amplification. Datadog hashes 4-grams while ES aggressively normalizes 3-grams.

Inside Husky’s query engine: Real-time access to 100 trillion events | Datadog

See how Husky enables interactive querying across 100 trillion events daily by combining caching, smart indexing, and query pruning.

datadoghq.com

Lucene is getting an increasing number of high-quality contributions from ByteDance employees, especially around performance. Good to see that this project keeps attracting contributors from all around the world.

Another common point I did not expect: Vespa's strict vs. unstrict iterators is quite similar to Lucene's two-phase iteration. And both projects use this feature to effectively combine dynamic pruning with filtering (a hard and underappreciated problem IMO).

Uwe now explains how Lucene takes advantage of the Panama foreign memory and vector support in spite of the fact that these features are still preview/incubating in the JDK

Bild

Yelp's nrtSearch was just upgraded to Lucene 10. Also switched from persistent storage to object storage as a source of truth, and plans on doing NRT replication via object storage instead of over the network. Very similar to Elasticsearch Serverless. engineeringblog.yelp.com/2025/05/nrts...

Nrtsearch 1.0.0: Incremental Backups, Lucene 10, and More

Nrtsearch 1.0.0: Incremental Backups, Lucene 10, and More Sarthak Nandi and Andrew Prudhomme May 8, 2025 It has been over 3 years since we published our Nrtsearch blog post and...

engineeringblog.yelp.com