AI Search and Recommendation Engine

We asked our AI search engine, “find me dresses and concerts.” It returned results, but somehow managed to miss both properly.

Our semantic search at AgileForce takes what a user types, embeds it, finds nearest matches, done. Simple, until someone types “dresses and concerts” – or just taps “yes” to a follow-up offering both.

We embedded that sentence and searched anyway. Results weren’t empty or obviously broken.

Just… mediocre.

Worse than either half alone would’ve been, which is the kind of failure that survives longest in production, because nobody can point at one specific bad result and say “that’s wrong.”

Turned out this wasn’t one bug, it was two, and only one was actually about vectors.

First, a single query vector can only represent one position in vector space. Ask for two fundamentally different things and you get one point sitting somewhere between the dress cluster and concert cluster – producing results that are vaguely both and properly neither.

Second, events and products need completely different filters. Events need things like date ranges, weekend flags and an event namespace. Products need audience classifications and a product namespace. Apply both namespaces in the same query and the conditions contradict each other, returning nothing.

Fixing the embedding wouldn’t have fixed the filtering. They had to be solved together.

So we stopped trying to answer one query with one search. Instead, one LLM call reads the message and splits it – figures out if it’s event, product, or both, extracts the relevant filters, and returns separate rewritten queries in a single structured response. One LLM call, not two. Each branch then runs its own search with its own filters.

The tricky part was the budget. A single search with 15 slots was already silently splitting itself between dresses and concerts through geometry alone – sometimes 12/3, sometimes 15/0, and you’d never know. Running two explicit searches makes that allocation visible, deliberate and tunable.

It’s still a guess. But at least it’s a guess we can inspect and adjust.

Last piece – scores from two different searches aren’t comparable. A 0.61 from the event search doesn’t mean the same thing as a 0.61 from the product search. So we don’t trust that sort. Everything gets pulled together and re-scored by one reranker, on one scale, against what the user actually asked.

Honestly, this is the kind of bug that teaches you more than the feature itself ever did.

If you’ve hit something similar, I’d genuinely love to hear how you solved it – drop a comment, or follow along, I’m sharing more of these as we build.

Share this :

Leave a Reply

Your email address will not be published. Required fields are marked *

Latest blog & articles

MING

A progress bar that just... froze. That's how this whole thing started. Another...

Sway

I believe the most dangerous performance bug is the one that looks fixed. ...

Get Free Consultation!