A customer types "something warm for hiking in the rain" into your store's search bar. Your keyword search returns nothing. No product in your catalog has those exact words in its title or description.
That customer leaves. Multiply this by hundreds of sessions a day.
I've been building search for e-commerce stores for the past year, and the pattern is always the same: keyword search works when customers know exactly what they want ("Nike Air Max 90 black size 42"). It fails when they describe what they need.
Your Search Engine Doesn't Know What "Gift for Dad" Means
Keyword search (Elasticsearch, Solr, PostgreSQL full-text) matches tokens. "Gift for a 5 year old boy" returns nothing in a toy store because no product title contains those words. Same with "laptop for video editing" when your catalog says "MacBook Pro M3 16GB RAM."
The customer describes a need. The search engine looks for exact words. Nobody finds anything.
Vector Search: The Opposite Problem
Vector search (embeddings) fixes the meaning problem. You encode products and queries into the same vector space, then find nearest neighbors. "Something warm for hiking" lands close to "insulated waterproof hiking jacket" because the embedding model understands meaning.
But vector search has its own problems:
SKU/model lookups fail. A customer types "XJ-4520" and vector search returns random products that happen to be close in embedding space.
Exact attribute matching is weak. "Red shoes size 38" might return blue shoes size 42 because the embedding thinks they're semantically similar (they're both shoes).
Precision drops with large catalogs. When you have 50,000+ products, the nearest neighbors might be "close enough" semantically but completely wrong for the customer.
Hybrid: BM25 + Vector Search
The solution I landed on: run both searches in parallel and merge results.
BM25 handles:
- Exact product names and SKUs
- Brand names ("Nike", "Bosch")
- Specific attributes ("size 38", "500ml", "red")
Vector search handles:
- Natural language descriptions ("something warm for winter")
- Intent-based queries ("gift for a coffee lover")
- Cross-language queries (customer asks in German, catalog is in English)
How the Merge Works
Both searches return scored results. The trick is normalizing scores so they're comparable, then combining them with configurable weights.
def hybrid_search(query: str, store_id: str, limit: int = 20):
# Run both searches in parallel
bm25_results = bm25_search(query, store_id, limit=limit * 2)
vector_results = vector_search(query, store_id, limit=limit * 2)
# Normalize scores to 0-1 range
bm25_scores = normalize(bm25_results)
vector_scores = normalize(vector_results)
# Merge with weights (tuned per use case)
merged = {}
for product_id, score in bm25_scores.items():
merged[product_id] = score * BM25_WEIGHT
for product_id, score in vector_scores.items():
if product_id in merged:
merged[product_id] += score * VECTOR_WEIGHT
else:
merged[product_id] = score * VECTOR_WEIGHT
return sorted(merged.items(), key=lambda x: x[1], reverse=True)[:limit]
The weights need tuning per store. Stores with lots of SKU-based lookups benefit from higher BM25 weight. Stores where customers describe what they want (fashion, home goods) benefit from higher vector weight. I don't have a universal formula. Start at 50/50 and adjust based on your query logs.
Cross-Encoder Reranking
After the merge, a cross-encoder reranker compares each candidate directly against the query.
Unlike bi-encoders (which encode query and product separately), cross-encoders take the pair as input and output a relevance score. More expensive, but more accurate.
from sentence_transformers import CrossEncoder
reranker = CrossEncoder("cross-encoder/ms-marco-MiniLM-L-6-v2")
def rerank(query: str, candidates: list[dict]) -> list[dict]:
pairs = [(query, c["text"]) for c in candidates]
scores = reranker.predict(pairs)
for i, candidate in enumerate(candidates):
candidate["rerank_score"] = float(scores[i])
return sorted(candidates, key=lambda x: x["rerank_score"], reverse=True)
I run this on the top 20-30 candidates from hybrid search, not the full catalog. This keeps response times reasonable since cross-encoders are slow on large sets.
Note: the code examples above are simplified for clarity. Production code needs error handling, async execution, and score caching.
Cross-Language Search
One side effect of using intfloat/multilingual-e5-large: it maps 100+ languages into the same vector space. A query in French against an English catalog returns correct results because the embedding model treats meaning, not language, as the proximity metric. No translation API needed. If you sell across borders, the multilingual embedding model does the work for free.
What This Doesn't Do
Limitations:
Image search. Customers can't upload a photo and find matching products. This is a different problem requiring CLIP or similar models.
Personalization. The search doesn't learn from individual user behavior. It treats every query independently.
Typo correction. Heavy typos can throw off both BM25 and vector search. I handle this with query preprocessing, but it's not perfect.
Real-time inventory. Search returns products that exist in the catalog. Stock availability is a separate check.
Stack
For anyone building something similar:
Vector DB: Qdrant (self-hosted). Fast, supports payload filtering, good Python client.
BM25: Qdrant's built-in BM25. No need for a separate Elasticsearch instance.
Embeddings:intfloat/multilingual-e5-large(1024 dimensions, 100+ languages)
Reranker:cross-encoder/ms-marco-MiniLM-L-6-v2
Orchestration: LangGraph for agent routing (product search vs support vs order tracking)
What I've Seen in Practice
I don't have clean A/B test data to share yet. What I can say from manually testing across several store catalogs:
- Keyword-only search fails on most natural language queries. If the customer doesn't use the exact product name, they get nothing.
- Vector-only search handles descriptions well but returns wrong results for SKU lookups and specific attributes (color, size).
- Hybrid search with reranking handles both query types. SKU searches still work. Descriptive queries return relevant products.
I won't put a percentage on it until I have proper metrics. If you're building this, set up evaluation before you ship.
I wrote about why I abandoned vector-only search after SKU lookups returned random products. The data sync pipeline that feeds the search engine is covered in Syncing 60,000 Products Without Breaking Everything.
I built this as part of Emporiqa, a chat assistant for e-commerce stores. Official Drupal module on drupal.org, WooCommerce plugin, Sylius plugin on Packagist, and a webhook API for anything else. You can test the search on your own catalog: the sandbox syncs up to 100 products in about 2 minutes. No credit card.
