Zum Hauptinhalt springen
Echtzeit-Radar & Feeds
Alle RSS Feeds ➔
👥 Community & Social
•••
IT Security NachrichtenBug in Microsoft Word: PDFs verschwinden beim Speichern spurlos(29.09.2026 um 08:43 Uhr)
•
IT Security NachrichtenGroupware Zimbra: Update schließt zahlreiche Sicherheitslücken(29.09.2026 um 08:31 Uhr)
•
IT Security NachrichtenGründer größter Drogenhandel-Plattform: „Bereue meine Taten“(29.09.2026 um 08:37 Uhr)
•
IT Security NachrichtenIT-Störung trifft Kommunalverwaltungen in Rheinland-Pfalz(29.09.2026 um 08:40 Uhr)
•••
IT NachrichtenMicrosoft updates Fabric for agentic transformation(29.09.2026 um 08:30 Uhr)
••••
IT Security NachrichtenBug in Microsoft Word: PDFs verschwinden beim Speichern spurlos(29.09.2026 um 08:43 Uhr)
•
IT Security NachrichtenGroupware Zimbra: Update schließt zahlreiche Sicherheitslücken(29.09.2026 um 08:31 Uhr)
•
IT Security NachrichtenGründer größter Drogenhandel-Plattform: „Bereue meine Taten“(29.09.2026 um 08:37 Uhr)
•
IT Security NachrichtenIT-Störung trifft Kommunalverwaltungen in Rheinland-Pfalz(29.09.2026 um 08:40 Uhr)
•••
IT NachrichtenMicrosoft updates Fabric for agentic transformation(29.09.2026 um 08:30 Uhr)
•
Intelligence View
⚡ tsecurity.de Intelligence

Hybrid Search for E-commerce: When Keywords Alone Fail

A customer types "something warm for hiking in the rain" into your store's search bar. Your keyword search returns nothing. No product in your catalog has those exact words in its title or description. That customer leaves. Multiply this…

0
↗ Quelle (dev.to)
Reagiere als Erste:r — dein Feedback zählt!

A customer types "something warm for hiking in the rain" into your store's search bar. Your keyword search returns nothing. No product in your catalog has those exact words in its title or description.



That customer leaves. Multiply this by hundreds of sessions a day.





I've been building search for e-commerce stores for the past year, and the pattern is always the same: keyword search works when customers know exactly what they want ("Nike Air Max 90 black size 42"). It fails when they describe what they need.






Your Search Engine Doesn't Know What "Gift for Dad" Means



Keyword search (Elasticsearch, Solr, PostgreSQL full-text) matches tokens. "Gift for a 5 year old boy" returns nothing in a toy store because no product title contains those words. Same with "laptop for video editing" when your catalog says "MacBook Pro M3 16GB RAM."



The customer describes a need. The search engine looks for exact words. Nobody finds anything.






Vector Search: The Opposite Problem



Vector search (embeddings) fixes the meaning problem. You encode products and queries into the same vector space, then find nearest neighbors. "Something warm for hiking" lands close to "insulated waterproof hiking jacket" because the embedding model understands meaning.



But vector search has its own problems:





  • SKU/model lookups fail. A customer types "XJ-4520" and vector search returns random products that happen to be close in embedding space.


  • Exact attribute matching is weak. "Red shoes size 38" might return blue shoes size 42 because the embedding thinks they're semantically similar (they're both shoes).


  • Precision drops with large catalogs. When you have 50,000+ products, the nearest neighbors might be "close enough" semantically but completely wrong for the customer.





The solution I landed on: run both searches in parallel and merge results.



BM25 handles:




  • Exact product names and SKUs

  • Brand names ("Nike", "Bosch")

  • Specific attributes ("size 38", "500ml", "red")



Vector search handles:




  • Natural language descriptions ("something warm for winter")

  • Intent-based queries ("gift for a coffee lover")

  • Cross-language queries (customer asks in German, catalog is in English)






How the Merge Works



Both searches return scored results. The trick is normalizing scores so they're comparable, then combining them with configurable weights.




def hybrid_search(query: str, store_id: str, limit: int = 20):
# Run both searches in parallel
bm25_results = bm25_search(query, store_id, limit=limit * 2)
vector_results = vector_search(query, store_id, limit=limit * 2)

# Normalize scores to 0-1 range
bm25_scores = normalize(bm25_results)
vector_scores = normalize(vector_results)

# Merge with weights (tuned per use case)
merged = {}
for product_id, score in bm25_scores.items():
merged[product_id] = score * BM25_WEIGHT

for product_id, score in vector_scores.items():
if product_id in merged:
merged[product_id] += score * VECTOR_WEIGHT
else:
merged[product_id] = score * VECTOR_WEIGHT

return sorted(merged.items(), key=lambda x: x[1], reverse=True)[:limit]






The weights need tuning per store. Stores with lots of SKU-based lookups benefit from higher BM25 weight. Stores where customers describe what they want (fashion, home goods) benefit from higher vector weight. I don't have a universal formula. Start at 50/50 and adjust based on your query logs.






Cross-Encoder Reranking



After the merge, a cross-encoder reranker compares each candidate directly against the query.



Unlike bi-encoders (which encode query and product separately), cross-encoders take the pair as input and output a relevance score. More expensive, but more accurate.




from sentence_transformers import CrossEncoder

reranker = CrossEncoder("cross-encoder/ms-marco-MiniLM-L-6-v2")

def rerank(query: str, candidates: list[dict]) -> list[dict]:
pairs = [(query, c["text"]) for c in candidates]
scores = reranker.predict(pairs)

for i, candidate in enumerate(candidates):
candidate["rerank_score"] = float(scores[i])

return sorted(candidates, key=lambda x: x["rerank_score"], reverse=True)






I run this on the top 20-30 candidates from hybrid search, not the full catalog. This keeps response times reasonable since cross-encoders are slow on large sets.



Note: the code examples above are simplified for clarity. Production code needs error handling, async execution, and score caching.





One side effect of using intfloat/multilingual-e5-large: it maps 100+ languages into the same vector space. A query in French against an English catalog returns correct results because the embedding model treats meaning, not language, as the proximity metric. No translation API needed. If you sell across borders, the multilingual embedding model does the work for free.






What This Doesn't Do



Limitations:





  • Image search. Customers can't upload a photo and find matching products. This is a different problem requiring CLIP or similar models.


  • Personalization. The search doesn't learn from individual user behavior. It treats every query independently.


  • Typo correction. Heavy typos can throw off both BM25 and vector search. I handle this with query preprocessing, but it's not perfect.


  • Real-time inventory. Search returns products that exist in the catalog. Stock availability is a separate check.






Stack



For anyone building something similar:





  • Vector DB: Qdrant (self-hosted). Fast, supports payload filtering, good Python client.


  • BM25: Qdrant's built-in BM25. No need for a separate Elasticsearch instance.


  • Embeddings: intfloat/multilingual-e5-large (1024 dimensions, 100+ languages)


  • Reranker: cross-encoder/ms-marco-MiniLM-L-6-v2


  • Orchestration: LangGraph for agent routing (product search vs support vs order tracking)






What I've Seen in Practice



I don't have clean A/B test data to share yet. What I can say from manually testing across several store catalogs:




  • Keyword-only search fails on most natural language queries. If the customer doesn't use the exact product name, they get nothing.

  • Vector-only search handles descriptions well but returns wrong results for SKU lookups and specific attributes (color, size).

  • Hybrid search with reranking handles both query types. SKU searches still work. Descriptive queries return relevant products.



I won't put a percentage on it until I have proper metrics. If you're building this, set up evaluation before you ship.






I wrote about why I abandoned vector-only search after SKU lookups returned random products. The data sync pipeline that feeds the search engine is covered in Syncing 60,000 Products Without Breaking Everything.



I built this as part of Emporiqa, a chat assistant for e-commerce stores. Official Drupal module on drupal.org, WooCommerce plugin, Sylius plugin on Packagist, and a webhook API for anything else. You can test the search on your own catalog: the sandbox syncs up to 100 products in about 2 minutes. No credit card.

Ähnliche Beiträge
🔍 Verwandte News

Auch interessante Nachrichten Hybrid Search for E-commerce: When Keywords Alone Fail

Thematisch verwandte Begriffe: Hybrid, Search, Ecommerce, When · 6 Treffer

Laden...

Beiträge werden geladen ...

Laden...

Videos werden geladen ...

Laden...

Beiträge werden geladen ...

Laden...

Videos werden geladen ...

Laden...

Beiträge werden geladen ...

Laden...

Videos werden geladen ...

Laden...

Beiträge werden geladen ...

Laden...

Videos werden geladen ...

💬 Kommentare werden geladen…
Zum Aktualisieren ziehen
ZERO-DAY CVE-2026-101281 | A flaw has been found in Trusted Domain Project OpenDMARC up to 1.4.2. …
Advisory →
tsecurity.de Icon
Offline-Lesen, Eilmeldungen & 0ms Ladezeit

Installiere tsecurity.de direkt auf deinen Home-Bildschirm für das ultimative Vollbild-Magazinerlebnis ohne Browser-Leisten.

Nächster Beitrag