back
Pons
Elixir/OTP · Python ONNX · Rust/Tantivy · RocksDB · FAISS · DVC · Cloudflare R2
status:active
ref:ML-01
size:—

Production hybrid search engine over 124M academic papers. Combines dense vector retrieval (FAISS) with BM25 sparse retrieval (Tantivy), fused via Reciprocal Rank Fusion. Self-hosted on a single VPS for ~€17/month, with the equivalent managed stack (Pinecone + OpenAI + Elasticsearch) costs $200–500/month.

The FAISS index is topic-sharded into 4,514 partitions with query routing via cosine similarity to topic centroids. This achieves p50 83ms, p95 110ms search latency at 9× throughput under 16 perfectly concurrent users. The embedding pipeline uses INT8-quantised ONNX models via Python port workers, supervised by an Elixir/OTP pool. This means no GPU dependency, CPU-efficient inference only.

The ranking metric is a novel normalised inverse Rao-Stirling diversity score. Standard Rao-Stirling measures cross-disciplinary citation breadth (how far a paper's references reach across fields.) I inverted and normalised it to produce a 0–1 score where higher values indicate papers that draw from more diverse disciplinary sources, surfacing underrepresented cross-disciplinary work alongside highly-cited papers within narrow fields.

The ML pipeline is fully reproducible via DVC and Cloudflare R2 across 5 stages: embeddings → FAISS shards → Tantivy index → RocksDB metadata store → lookup tables. Everything is re-runnable from raw data. The metadata store holds 124M keys in ~4.7GB via msgpack-compressed RocksDB, enabling O(1) lookups at query time with no database round-trip.

Embedding model selection was evaluated across MiniLM and Mixedbread variants using MTEB benchmarks. All data maintained by Pons is GDPR compliant as we removed all PII from the set.

— —