MindzKonnected – site header (partial)

Author: Shikhar

  • Grade the Answer, Not the Query: How We Built MindzEval to Evaluate Our Text-to-SQL Agent

    We built an agent that turns plain-English questions into SQL, and it worked. Then someone asked the one question it could not answer for us: how do we know it is right? This post is about the evaluation suite we built to replace a feeling with a number.

    If you ask our agent “how many sci-fi films are there?”, it works out what you mean, maps your words to the right database fields, writes the query, runs it, and hands back the answer. Throughout this post, the examples use a film catalogue as a stand-in for our real data.

    However, every time we changed a prompt, swapped the model, or edited the semantic layer, we were guessing. We could not prove we had not made the agent worse, so if accuracy quietly dropped, we would find out from a user weeks later. What we wanted was one number we could check ourselves, which meant the same thing every week and told us where the agent broke when it fell.

    Text-to-SQL evaluation at a glance: six problems and fixes, plus a first baseline of about 7% execution accuracy

    Why doesn’t “it looks right” work for testing an agent?

    Before the suite existed, checking the agent meant typing in a few questions and reading the answers. That approach fails in three ways:

    • It cannot be repeated, so you can never say “this is better than last time” and mean it precisely.
    • It produces prose rather than a number, and “seems good” cannot go on a graph or block a bad change from merging.
    • Its coverage is limited to whatever you thought of, even though the questions that break an agent are usually the ones nobody thought to type.

    In other words, we needed something objective, reproducible, and diagnostic.

    Why not just use RAGAS or DeepEval?

    Two tools stood out:

    • DeepEval, an open-source framework for testing LLM applications with pytest-style assertions and a large library of ready-made metrics.
    • RAGAS, whose text-to-SQL evaluation guide runs the known-correct SQL and the agent’s SQL, compares the data each returns, and improves the prompt through error analysis. Its worked example climbs from about 2% accuracy to roughly 71% across three prompt versions.

    That loop of measuring, analysing, fixing, and measuring again became the backbone of everything below, and we took it wholesale.

    What the tools could not give us

    Even so, neither tool fitted our agent on its own, because we needed two things they could not provide:

    • We needed to grade the stages, not just the ends. Our agent is a pipeline: intent detection, then a semantic layer that maps business words to database fields, then a query planner and a SQL compiler, and only then execution. Off-the-shelf tools grade what goes in and what comes out, so they cannot see the semantic layer choosing the wrong mapping three stages before any SQL exists.
    • We needed to score the right to refuse. If a question is hopelessly vague, the agent should ask you to narrow it down, and if the data cannot answer it, the agent should say so instead of inventing a number. A tool that only compares result sets cannot express “the right move here was to refuse”.

    AI agent evaluation across the text-to-SQL pipeline: intent detection, semantic layer, query planner, SQL compiler, execution

    Introducing MindzEval

    MindzEval is our evaluation suite for the text-to-SQL agent. It keeps the RAGAS loop of measuring, analysing, and fixing, and adds what the tools could not give us:

    • It runs every test question through the full pipeline and grades each stage, not only the final answer.
    • It scores refusals alongside results, so a correct “please narrow this down” counts as a success.
    • It reports one number you can trust, broken down into named causes whenever that number falls.

    How MindzEval works: six design decisions

    Each decision below started as something that went wrong and ended as a rule MindzEval now follows. Decisions 1 to 3 make the score correct and useful, Decisions 4 and 5 cover our two needs, and Decision 6 makes sure MindzEval itself can be trusted.

    Decision 1: Compare results, not SQL text

    Why it matters

    Checking the agent’s SQL against ours letter by letter does not survive contact with reality, because the same query can be written many correct ways:

    • a different join order
    • IN instead of a subquery
    • a CTE instead of a nested SELECT

    All of them return identical data, yet all of them are “wrong” if you compare the text.

    What we did

    • We grade the answer, not the query. We run the known-correct query and the agent’s query, and check whether the two return the same films. The share of questions where they match is what we call execution accuracy.
    • Think of a maths exam. It is like marking by the final answer rather than insisting the working matches yours line for line: the returned films are the answer, and the SQL is the working.
    • We compare on identity, not whole rows. The agent returns different columns for different questions, sometimes just a title and sometimes a title, a release year, and a box office total. Every film has a stable id, film_id, in every row, so we compare the set of ids and let the columns vary freely. That asks the only question that matters: did it find the right films?

    Execution accuracy for LLM-generated SQL: compare query results, not SQL text, since different queries return the same rows

    Because “the same answer” differs by question type, the comparison has four shapes:

    Question type Example What we compare
    Filter “documentaries” The set of film ids
    Ranking “top 10 by box office” The ordered list of ids, with the set checked separately so we can tell “wrong films” from “right films, wrong order”
    Count “how many sci-fi films are there?” A single number
    Per-genre breakdown films grouped by genre The group totals

    Decision 2: Store the query, not the rows

    Why it matters

    The obvious approach is to save the list of films each question should return. However, the moment new data loads, every saved list goes out of date, and the suite fails on cases where the agent is perfectly fine. Consequently, you spend your mornings re-recording answers instead of fixing the agent.

    What we did

    • We store the known-correct SQL and run it fresh on every run, so the reference answers stay correct when the data refreshes.
    • The golden answer is the query, not the number. For “films from Spain”, it is not “2,662 films” but the query that counts them, whatever today’s number is.
    • Each case also records what each stage should produce, in its expect block, which is how MindzEval grades individual stages rather than only the output.

    Golden SQL vs golden rows in an LLM evaluation dataset: rerunning golden queries prevents false failures after data refresh

    A case ends up looking like this:

    - id: filter_documentary_001
      question: "Which documentaries are in the catalogue?"
      category: filter
      difficulty: easy
      golden_sql: |
        SELECT DISTINCT fg.film_id
        FROM public.v_film_genres fg
        JOIN public.v_genres g ON g.genre_id = fg.genre_id
        WHERE g.genre_label = 'Documentary';
      compare_on: film_id_set
      expect:
        semantic: {status: RESOLVED}
        plan: {operation_type: FILTER}
        terminal_status: SUCCESS

    There is a related trap: golden queries must use the values actually in the database.

    • The data says 'España', not “Spain”, and 'Sci-Fi & Fantasy', not “sci-fi”.
    • A question about wildlife and nature films once came back empty for exactly this reason, whereas it now returns the 37 films it should.

    As a result, a good chunk of the dataset is quietly a set of regression tests for bugs that actually happened.

    Decision 3: Break one number into several

    Why it matters

    “The agent is 7% accurate” is true but useless on its own. It does not say whether queries ran and returned the wrong thing or never ran at all, nor whether filters are too loose or too strict, so you cannot act on it.

    What we did

    • Executable rate answers “did the query even run?”, which separates a wrong answer from a broken one.
    • Precision and recall on the film sets show the direction of the error. High recall with low precision means the filter drags in extra films, while the reverse means it is too strict.
    • Failure buckets sort every failing case into exactly one named cause, because a failure with a name is a failure you can fix.
    • Slicing by question type and difficulty turns one headline figure into a sentence such as “94% on simple filters, 45% on multi-constraint aggregations” (an illustration of the format, not our current result), which actually decides what you work on next.
    Failure bucket What it means
    Query never ran The agent did not execute any SQL
    Asked for clarification The agent asked the user to narrow the question
    Fell back The agent took its fallback path instead of answering
    Ran out of step budget The agent hit its limit before finishing
    SQL error The query was executed but threw an error
    Confident answer, no query (hallucination path) The agent returned a successful-looking answer without running any query

    Text-to-SQL agent error analysis: executable rate, precision and recall, and failure buckets that expose AI hallucinations

    The last bucket matters most, because in a text-to-SQL system a number with no source behind it is the most dangerous output there is.

    Decision 4: Score refusals both ways

    Why it matters

    This is the second need from earlier: sometimes the right answer is no answer.

    • “Show me the good films” is too vague, so the agent should ask what “good” means.
    • “What is the weather” is out of scope, so it should decline rather than guess.

    A metric built on comparing rows cannot score these cases, and leaving them out creates a blind spot exactly where a confident wrong answer does the most damage.

    What we did

    • A separate behavioural score checks whether the agent asked for clarification, fell back gracefully, or flagged a concept as unsupported when it should have.
    • It scores both directions. An agent that asked for clarification on every question would score perfectly and yet be useless.
    • Negative controls catch over-refusal. These are questions the agent should simply answer, and we score them too, because knowing when not to refuse is as much a skill as knowing when to refuse.

    Evaluating AI agent refusals: correct answers, over-refusal, confident wrong answers (hallucinations), and correct clarification

    Decision 5: Let one file know the agent

    Why it matters

    Grading every stage, our first need, means reading the agent’s intermediate results: what it understood, the plan it built, and the SQL it compiled. However, these are buried in its message log, and digging them out depends on exact tool names and message shapes. Since the agent is under active development, that knowledge goes stale the instant someone refactors.

    What we did

    • Every metric reads from a normalised turn record, a clean record of one turn that contains no agent-specific types.
    • Exactly one module builds that record, and it is the only file in MindzEval allowed to import the agent or name a tool. As a result, a refactor breaks one file instead of every metric.
    • One run serves every metric, so we never pay to re-run the agent for each one.
    • A small test guards that file by checking that it still recognises the agent’s real tool names.

    Although that test sounds trivial, it caught a real bug: code looking for a tool called semantic_resolver, which does not exist, had left an entire stage silently empty.

    Decision 6: Make the suite check itself

    Why it matters

    A measuring instrument can be broken too. If the extractor silently returned nothing, every metric would read zero, and it would look as though the agent had got much worse overnight. A number you cannot trust is worse than no number, because you will act on it with confidence.

    What we did

    • Two counts top every report: failures caused by bugs in our own harness, and failures caused by an empty extractor. Both are currently zero, which is how we know the score is real.
    • Crashes are attributed, so an agent bug is filed as a finding and our bug is filed as ours.
    • The harness has three dozen tests of its own, because the ruler needs its own ruler.

    LLM evaluation harness design: one extractor isolates agent internals so metrics survive refactors, with regression tests

    What did the first baseline tell us?

    When we ran MindzEval, the first full baseline came back at about 7% execution accuracy. We read that as good news, because a low number that breaks into named, fixable causes means the instrument is measuring something real. After all, the number’s job is not to be high on day one but to move and to say why when it does not.

    Finding Cases What it means
    Truncated, but everything returned was correct. 30 The 50-row display cap made a full match impossible even though the agent’s logic was right.
    Answerable questions not answered from a real query 35 The agent fell back or asked to clarify instead of just querying
    Of which: confident “success” with no query run 11 The hallucination path

    Two structural issues, rather than reasoning failures, dominated the score:

    • A display setting. The agent caps results at 50 rows, but our reference answer for “sci-fi films” is 215. It returned 50 correct films, so the cap alone made a full match impossible. MindzEval reports this as its own category, and it accounted for 30 cases.
    • The agent holding back. On 35 answerable questions, the agent did not answer from a real query, and it mostly fell back or asked for clarification. However, 11 of those 35 returned a confident “success” without running any query, which we would never have spotted by reading answers.

    In other words, instead of “the agent is 93% broken”, we got a to-do list.

    How we run it day to day

    • On every change, we run a fast slice of ten cases, which takes about seven minutes.
    • Every night, we run the full set of one hundred cases.
    • Later, DeepEval comes back in, both for wiring MindzEval into CI and for an optional judge that checks whether the final wording invents anything the data does not support.

    The core, however, stays ours, because no framework can see inside our agent or say that refusing was the right call.

    The lesson

    Grade the thing that matters, not the thing that is easy to compare. Matching SQL text is precise but wrong, whereas checking whether the agent found the right films is the only number worth trusting. A number you can check yourself, even an ugly one, beats a confident feeling every time.

    Want to know whether your agent is right, not just whether it looks right?

    MindzEval grades every stage of a text-to-SQL pipeline, scores refusals in both directions, and checks its own results on every run. To find out what it could reveal about your agent, call us on +91 85957 97775.

  • From ReAct to Deep Agent: When One Routine Is Not Enough

    Our chatbot had one way of thinking, and it used that one way for everything. This post is about the deep agent architecture we built to replace it.

    Ask it “hi” and it would stop, reason about what it needed to know, decide to search, run the search, read the results, and reason again about whether to search once more. Ask it to compare five vendors on price, latency, and compliance and it would do exactly the same thing. Same loop, same tools, same number of sources pulled back. The only difference between those two requests was how many times the loop spun before the model decided it had enough.

    We shipped it anyway, because it worked. People used it. But as the questions got harder, we kept hitting the same problem: slow on the easy things, shallow on the hard ones, and equally confident either way.

    The problem was never that the model was not smart enough. The problem was that our architecture handed it exactly one routine and no way to choose a different one. A greeting and a research assignment went through identical machinery, and the machinery could not tell them apart.

    This is the story of replacing that single routine with a deep agent: an architecture that first decides how much effort a question deserves, then brings only the machinery that effort actually requires.

    The original design: one ReAct loop for everything

    Our chatbot brain ran a ReAct loop. ReAct is short for Reason and Act. The model reasons about what it needs, takes an action (usually a tool call), observes what comes back, and reasons again. It repeats until it decides it can answer.

    Think of hiring a research assistant. You ask a question. They pause and think about what they need to know. Then they go look something up and read it. They ask themselves whether they have enough, decide, and repeat.

    A greeting, a quick lookup and a five-way comparison all passing through the same Reason, Act, Observe loop at identical cost

    ReAct became the default for good reasons. It is a few dozen lines of code. It also handles open-ended questions without anyone having to anticipate them in advance. And it degrades gracefully, because a model that does not need a tool simply answers.

    But a single ReAct loop has four structural properties that only reveal themselves as problems once real traffic hits it.

    • It has one routine. Nothing in the loop ever asks how hard the incoming question is. Every request enters the same machinery at the same depth.
    • It trusts its tools. Whatever a tool returns goes straight into context as fact. There is no gate between retrieval and reasoning.
    • It has one shared context. Every observation from every step piles into the same message history, and the final answer is written from that same pile.
    • It has no plan. The model may reason, but nothing obliges it to, and nothing tracks which parts of a question are still unanswered.

    What a deep agent actually is

    A deep agent is not a smarter model. It is a harness built around the same tool-calling loop, with scaffolding that handles the things a bare loop handles badly. LangChain’s deepagents library is the reference implementation, and it organises that scaffolding into four groups.

    Deep agent architecture: an unchanged tool-calling core wrapped in execution environment, context management, delegation and steering

    • Execution environment. The agent gets tools, but also a virtual filesystem it can read and write, permission rules over which paths it may touch, and optional sandboxed code execution. Work can live outside the context window, in files.
    • Context management. Skills load domain knowledge on demand rather than upfront. Memory files persist preferences and conventions across sessions. Summarization and offloading compress long histories and oversized tool results automatically. On supported models, the static parts of the prompt are cached.
    • Delegation. The agent can maintain a structured task list as it works, and it can spawn subagents. A subagent runs in its own fresh context, works autonomously, and hands back a single distilled report.
    • Steering. Sensitive tool calls can pause and wait for a human to approve, edit, or reject them before they run.

    The unifying idea is simple. In short a plain ReAct loop has one context, one routine, and one level of effort. A deep agent has several of each, plus the ability to decide which one a given request should get.

    We did not adopt all of it. We took the capabilities that mapped onto problems we actually had, and where the harness had no answer, we built our own on top of it, problem by problem, some fixes came free with the harness and some we added ourselves. Each section below says which is which.

    Problem 1: Every question got the heavyweight treatment

    The problem

    A greeting should be instant. A factual lookup should be quick. A five-way comparison should be thorough. We gave all three the same routine: reason, search, observe, reason again.

    The cost ran in both directions. Simple questions took seconds longer than they needed to and burned tokens on research that was never necessary. Worse still, the loop’s bias toward acting meant that on easy questions the agent would go looking things up that it already knew, and come back with a weaker answer than if it had simply spoken.

    The deep agent fix: triage before effort

    Before any real work happens, one fast classification pass reads the question and sorts it into a lane.

    • DIRECT. Greetings, follow-ups, clarifications. Answer from what the model already knows. No tools at all.
    • TOOL. One clear factual question, like the capital of France or the height of the Eiffel Tower. One lookup which runs through the web search tool we built ourselves, one answer, stop.
    • DEEP_RESEARCH. Genuinely open-ended work. Compare these approaches, trace the history of this, investigate this properly. This lane gets the full loop.

    class Intent(str, Enum):

        DIRECT        = “DIRECT”          # no lookups needed

        TOOL          = “TOOL”            # one lookup

        DEEP_RESEARCH = “DEEP_RESEARCH”   # full investigation

    Each lane only assembles the machinery it needs. Simple messages skip the heavy path and come back faster and cheaper. Hard questions still get the full investigation. They just stop subsidising the easy ones.

    One rule governs the whole thing: when triage is unsure, it picks the most thorough lane. If we cannot tell how hard a question is, we spend more to be safe. Saving tokens never outranks being right.

    This is a piece we built ourselves. The harness gives you the machinery to run work at different depths. Deciding which depth a particular question deserves is application logic, and you have to write it.

    One triage pass sorting an incoming message into three lanes: DIRECT with no lookups, TOOL with one lookup, DEEP_RESEARCH with a wide sweep

    Problem 2: Bad sources went straight into the answer

    The problem

    When the agent looked something up, whatever came back went directly into its context. If a search returned four articles and one was only loosely related, all four arrived carrying equal authority. The model had no way to separate a strong source from a weak one, so it treated them alike and answered confidently from the mixture.

    This failure is worth naming precisely, because it usually gets mislabelled. The agent was not hallucinating. Instead, it was faithfully repeating what retrieval handed it. The fault was upstream of the model entirely.

    The deep agent fix: a grader between the tool and the loop

    We put a checking step between retrieval and reasoning. Every time the agent looks something up, that step grades how relevant the results actually are before any of them are allowed into context.

    Good sources pass through. If they are bad, the system rewrites the query and searches again from a different angle. When the second attempt also comes up short, it flags the gap so the agent answers cautiously instead of confidently. Mixed results get split: it keeps the good ones and searches again to fill what is missing.

    The important detail is that the grader does not know what answer the agent is hoping for. As a result, it judges sources on their merits alone, which means it cannot talk itself into approving weak research just because the weak research would be convenient.

    This is the second piece we built. A deep agent harness controls what enters context and when. Judging whether a specific retrieval result is good enough for your domain is a call only you can make. This is the same instinct we applied one level down, when we rebuilt our crawler to return only the parts of a page that match the query.

    A grader scoring search results before the assistant sees them, routing to four outcomes: pass, refill, retry, or answer with caution

    Problem 3: The agent searched instead of thinking

    The problem

    The original loop could reason. Nothing made it. Handed a hard question, it would fall into a searching rhythm: query, read, query, read, never pausing to ask whether the last result actually helped or what was still missing.

    Multi-part questions exposed this most clearly. Asked four things at once, the agent would answer two well, mention a third in passing, and quietly drop the fourth. Nothing anywhere in the architecture was keeping track of the fact that a fourth part existed.

    The deep agent fix: make planning a step, not a suggestion

    This is where a deep agent’s task planning earns its place. The harness offers a task list the agent maintains as it works, with every item tracked as pending, in progress, or complete. A four-part question becomes four tracked items, and none of them can go missing without it being visible.

    We paired that with a forced rhythm: reflect before searching, and reflect again after every result before deciding the next move.

    It sounds trivial. Forcing the rhythm is exactly what turns frantic searching into deliberate investigation. It also brought a benefit we did not expect to value as much as we now do. Because the reasoning is written down as structured state, so when an answer goes wrong we can see precisely where the thinking went wrong instead of guessing at it.

    Before and after: unbroken searching that quietly drops part of a multi-part question, versus reflect-search-reflect with every part covered

    Problem 4: Long investigations buried their own findings

    The problem

    A hard question needs many lookups. In the original design, the result of every lookup piled into one shared context, and none of it ever left.

    The effect was perverse: answers got worse the longer the agent worked. Early findings, often the most relevant ones, ended up buried under later raw material. By the time the agent sat down to write, the good material from step two was competing for attention with fifteen pages of noise from step nine. In other words, more effort, worse result.

    The deep agent fix: delegate to subagents with their own context

    This is the capability that made deep agents worth adopting rather than just patching the loop.

    For the heaviest questions, the work now splits between a manager and a specialist. The manager holds the question and the plan. It delegates the actual searching to a subagent that runs in a completely separate context, works through its subtask alone, and hands back only the distilled findings.

    The manager never sees the raw pile. As a result, its context stays small and focused on what it is actually responsible for, which is the shape of the final answer. The subagent can go as deep as it needs to, because all of its clutter dies with it.

    Alongside that, summarization and offloading run automatically. History gets compressed as it ages, and oversized tool results are moved out of context and into files the agent can read back on demand.

    This part is native. Isolated subagent context, automatic summarization, and offloading to a filesystem are exactly what the harness is built to provide.

    Delegation in the deep agent architecture: a manager holds the question and plan while a specialist searches in its own workspace and returns only distilled findings

    Problem 5: Every lookup returned the same amount of material

    The problem

    Retrieval was configured once, globally. Every search pulled back the same fixed number of sources. That number was simultaneously too small for “list everything we know about X” and far too large for “what time does the gate open.”

    The deep agent fix: scale depth to the intent we already know

    We were already classifying every question in Problem 1, and that classification carries real information about breadth. So we let it set retrieval depth as well. Pinpoint questions get a small, precise set. Broad questions get a wide sweep.

    Same retrieval machinery, effort matched to the need, and no second classification pass to pay for.

    What changed

    The original design was not broken. It was undiscerning. It spent the same effort and ran the same strategy for every request, regardless of what the request actually was. Every fix we made added a piece of judgment.

    How the deep agent architecture changed five things: triage, a fact checker, required reflection, a manager and specialist split, and depth scaled to the question

    The part we are happiest about is what did not change. The original ReAct agent still exists in our system, untouched. The deep agent slots in as the default and calls the old loop when the old loop is the right tool for the job. We swapped the brain without rewiring anything around it.

  • Returning Only What Matters: Smarter Web Crawling with Semantic Search

    Two questions about the World Cup 2026 Final compared: a top-scorer question answered by the search preview alone, and a minute-by-minute breakdown that requires crawling the full page.

    When our agent needed to answer a complex question, it could not rely on search results alone. Sometimes the answer was buried deeper in a webpage. For eg: “Give me minute by minute breakdown of the world cup 2026 Final game” requires going deeper into the content of a related article as compared to “Who was the top scorer for Spain in the world cup 2026 Final?” which could be found in search results alone. 

    We covered how our web search tool found relevant URLs in our previous blog.

    In this blog, we are going to cover how our web crawl tool finds the relevant content inside a URL through semantic search.

    When search results are not enough

    A user would ask a question, and web_search would return five URLs that looked relevant. The agent would check the first result, get the description from the preview, and give an answer based on that.

    Often the answer was right. A simple question like “how many goals did Ronaldo score against Croatia in World Cup 2026?” could be answered from the preview alone. The search result would have the number right there in the description or first line of the page. Done.

    Sometimes it was not enough. A more complex question like “give me a minute by minute breakdown of the events in the Portugal vs Croatia game” could not be answered from the preview. The preview might just say “exciting match with 3 goals” or something generic. The real breakdown of what happened at each minute was deeper in the page, buried in a full match report or article. We had to crawl the entire URL to find the detailed information.

    When that happened, we needed to go deeper. We needed a second tool called web_crawl. This tool takes a URL and reads the full page, looking for the specific information the user asked for.

    That is when the second problem started.

    The leftover problem

    Our early version of web_crawl did what we asked it to do. It took a URL and returned the entire content of that page. All of it. Every word.

    If a user asked “what is the return policy” and we had to crawl a product page, web_crawl would return not just the return policy, but the product description, customer reviews, navigation menu, footer, everything that was on the page.

    Or if a user asked “how many goals did Ronaldo score against Croatia” and we crawled a sports news article, it would return the entire article, including player biographies, team history, match statistics for other games, and commentary. All of it, even though the answer was just one number.

    This worked, technically. But it created two real problems.

    A crawled match report split into eleven regions, only one holding the answer: about 2,000 of the 200,000 characters returned were relevant, roughly one chunk in a hundred.

    Why that hurt

    First, all that extra content wasted the model’s context tokens. If a user asked “how many goals did Ronaldo score against Croatia” and we crawled a full sports article with 5000 words of match details, player stats, and team history, we were burning tokens on 4999 words of garbage to answer a question that needed one number. With many queries running, those wasted tokens added up fast and made everything slower and more expensive.

    Second, dumping unrelated information confused the model. When a model reads a page full of product reviews mixed with the return policy mixed with shipping information, it gets confused about what matters. It might pull information from the wrong part of the page, or make up an answer based on the noise. The real answer gets buried.

    It is like asking a librarian for one specific fact and having them hand you the entire book instead of just the page you need. You have to read through all the noise to find the answer, and you might miss it or get lost in the details.

    A full crawled page using about 50,000 tokens of a 128,000-token context window versus about 5,000 after filtering, and unrelated passages leading the model to answer from the wrong match.

    The shift in thinking

    We realized that the problem was not with how we were getting the page. It was with what we were keeping from it.

    The goal was not to return everything. The goal was to return only what the user asked for. So instead of giving back the whole page when web_crawl ran, we needed a way to figure out which parts of the page actually matched the meaning of the user’s query.

    This was different from just searching for words. If a user asked “minute by minute breakdown of the Ronaldo vs Croatia game,” a simple word search would find pages with those words in them. But it would not know which parts of the page had the actual breakdown and which parts had other information like team stats or historical context. We needed web_crawl to understand meaning, not just match keywords.

    Keyword matching ranks chunks about ticketing, broadcasting and stadium costs highest, while semantic matching finds the minute-by-minute breakdown that shares no words with the query. 

    The fix

    We used a small AI model called sentence-transformers/all-MiniLM-L6-v2.

    Here is how it works. The model takes the user’s query and converts it into a set of numbers that capture the meaning of that query. Then it does the same thing for each chunk of the page. It breaks the page into smaller pieces and converts each one into numbers that capture its meaning.

    The magic is that these numbers put similar meanings close together. So “return policy” and “how to send items back” end up close to each other in this space, even though the words are different. “free shipping” and “no shipping cost” also end up close together. And “minute by minute breakdown” and “events that happened during the match” are also nearby in meaning, even though they use different words.

    Cosine similarity with all-MiniLM-L6-v2: each chunk becomes a 384-dimensional unit vector, scored against the query from 0.74 down to 0.03, with the ten closest chunks kept.

    Once we have these numbers for the query and all the chunks of the page, we find which chunks are closest to the query. Those are the chunks that mean the same thing as what the user asked for. We keep only those chunks and return them to the model.

    Eight-step pipeline turning a 200,000-character page into ten ranked chunks: strip links, split into 2,000-character chunks with 200 overlap, check a Redis embedding cache, embed, score, and keep the top ten.

    The function kept the same inputs. It still took a query and a URL. It just got smarter about what it gave back.

    Architecture of our web crawl pipeline: chat backend, MCP API, Redis streams, a Crawl4AI crawler worker, and a semantic worker that narrows 200,000 characters to about 20,000 before results return to the agent.

    The result 

    Before and after for the same URL and question: the whole page at about 50,000 tokens versus ten ranked chunks at about 5,000 tokens, a 90% reduction.

    Now when a user asks about a return policy, web_crawl gives back only the return policy section. When they ask about a minute by minute breakdown of the Ronaldo game, it returns only the breakdown section from the article, not the player biographies or team history. When they ask about shipping, it returns the shipping information. No more entire pages full of unrelated content.

    Lessons Learned

    This fixed both problems. The context is clean and focused, so we do not waste tokens on garbage. The model reads only what it needs and does not get confused by unrelated text. There are fewer hallucinations because there is less noise to confuse the model. The results are faster and cheaper because we use fewer tokens.

    The real lesson is this. Our web_search tool gets us started with quick answers. But when we need to go deeper into a page, web_crawl now does it smart. It does not just dump the whole page. It finds only what matters.

    Matching by meaning beats matching by volume. Giving less, but more relevant, information is better than giving everything. In almost every case, thirty words on topic are better than a thousand words scattered all over the place.

  • Crawl4AI vs SearXNG: Choosing the Right Web Search Tool for AI Agents

    At a glance: Our AI agent’s first web search tool used Crawl4AI, which opened a new headless Chromium browser for every search. Since each browser uses an estimated few hundred MB of memory, a load test with 10 concurrent searches exhausted the server’s memory and crashed it. We rebuilt web search for AI agents with SearXNG as the default, a plain HTTP metasearch call with no browser, and kept Crawl4AI only as a fallback that now runs one shared browser with capped concurrency. The lesson: the most powerful tool is not always the right default.

    Our AI agent needed to search the web. At first, this sounded like a solved problem, so we reached for a popular, powerful tool: Crawl4AI, which reads pages by driving a real headless browser. In fact, it worked perfectly for a single search. However, when we load-tested it with 10 searches running at the same time, the server ran out of memory and crashed. In this post, we explain why that happened and why the tool itself was not the problem. Finally, we show how we rebuilt web search for AI agents around a much lighter default, SearXNG, while still keeping Crawl4AI for the pages that truly need a browser.

    The first attempt

    Our first version used a tool called Crawl4AI. Crawl4AI is an open-source Python library built for collecting web content for AI systems. The way it works is that it opens a full, real web browser in the background, a headless Chromium browser, and uses it to load and read a page the same way a normal browser would. It is a popular and widely used library, with more than 60,000 stars on GitHub, so it felt like a solid, well supported choice.

    The crash

    Why Crawl4AI handles hard pages so well

    Because Crawl4AI drives a real browser, it can do things a plain request cannot. It can wait for the page to finish loading, scroll down to pull in content that only appears as you go, run the page’s own scripts, and then hand back a clean, readable version of the page. This is why it works so well on difficult websites, the ones that build their content with JavaScript after the page first loads, where a plain request would often return only an empty shell. That same power is also where our problem started.

    How much memory does a headless browser use?

    A real browser is heavy. Running headless Chromium is close to running a full copy of Chrome, with all of its moving parts loaded into memory at the same time. As a result, every browser instance that Crawl4AI opened used a large amount of memory. Public estimates put a headless Chromium instance at roughly 100–500 MB, with 300–500 MB a common planning figure once it renders a real page.

    What happened under load testing

    On top of that, our early script opened a brand new browser every single time it ran a search, instead of reusing one. So when many searches ran together, we were not opening one browser and sharing it. We were opening a separate browser for each search. Ten searches at the same time meant ten separate browsers. Each one carried its own few hundred MB, and none of that memory was shared between them, so the total added up fast with every extra search running at once.

    Two other things made this worse. First, when you run Crawl4AI on your own server, you have to manage all of this yourself, including the browsers and how many of them run at once. Second, the memory cost does not go down as you do more work. Every extra page you want to read at the same time means another full browser, and another few hundred MB on top.

    Reading one page at a time was fine. The trouble started the moment we needed to handle many searches at once, which is exactly what a real product has to do. We saw this clearly when we used JMeter to put more load on the application. When we reached 10 concurrent searches, the many browsers running at the same time filled up the server memory, and the server crashed.

    Diagram showing how 10 concurrent web searches each opened a separate headless Chromium browser using 150–400 MB, together using 1.5–4 GB of memory and crashing the server.

    The realisation

    When we looked at what had happened, we saw the real mistake. We had reached for the most powerful tool and made it our default. But most of our searches were ordinary. They did not need a full browser to read and render an entire page. They only needed a quick and light way to get search results. The powerful tool was not wrong. It was just the wrong choice for the common case.

    The fix

    We changed our default to a tool called SearXNG.

    SearXNG is a free and open-source metasearch engine. A metasearch engine does not keep its own index of the web. Instead, it sends the query to several other search engines, collects their results, and combines them into a single list. Because of this, it does not need to open a browser at all. It talks to the search engines directly and returns results, which is fast and light on memory. It is written in Python, it is released under the AGPL-3.0 open-source license, and it can be self hosted, which means we run it on our own server and stay in control of it. It also does not track or profile its users. It can pull results from many sources, including Google, Bing, Brave, DuckDuckGo, Qwant, Startpage, and Yahoo.

    Right now our default setup uses SearXNG together with DuckDuckGo. SearXNG is the layer that runs the search and gathers the results, and DuckDuckGo is one of the search engines it draws those results from. This combination gives us clean search results without the heavy cost of a browser.

    We did not throw Crawl4AI away. Some pages really do need a full browser to be read properly. So we kept Crawl4AI as a backup, ready to step in for those harder pages, instead of using it for every search.

    Crawl4AI vs SearXNG at a glance
    Crawl4AI SearXNG
    What it is Open-source Python crawler that drives a headless browser Open-source, self-hosted metasearch engine
    Uses a browser Yes, headless Chromium No, a plain HTTP call that returns JSON
    JavaScript-heavy pages Handles them well Returns search results, not rendered pages
    Memory per request Public estimates of roughly 100–500 MB per browser Far lower, since no browser runs
    Best for Hard pages that need full rendering Everyday searches
    How we use it Fallback only, with one shared browser Default for every search

    Architecture diagram of our web search pipeline: requests go through an MCP API and a Redis job queue to a crawler worker that uses SearXNG by default and Crawl4AI only as a fallback.

    How our web search works now

    Our web search now runs as a small job pipeline instead of a single function call. First, the chat backend sends a request to our MCP API at /api/web-search. The API verifies the user’s JWT, checks that they have access to the web search tool, and then creates a job ID.

    Next, the job goes into Redis. Its state is stored under its own key, and the work itself is pushed onto a Redis stream. At the same time, the backend subscribes to a Pub/Sub completion channel for that job before it starts waiting. As a result, the backend never polls for results; instead, it is notified the moment the job finishes.

    A crawler worker then claims the job from the stream and decides what kind of work it is: a search, or a page to read. For searches, it uses SearXNG, which is a plain HTTP call that returns JSON, with no browser involved. Only when SearXNG fails, or a page needs JavaScript to render, does the worker fall back to Crawl4AI. Even then, it no longer opens a new browser per request. Instead, all fallback jobs share one headless Chromium browser with capped concurrency, which fixes the exact mistake that crashed our first version.

    Finally, the results are stored back in Redis. Search results skip our semantic re-ranking step, because the search engines have already ranked the snippets. (We cover semantic re-ranking for full pages in our post on smarter web crawling with semantic search.) A response aggregator then builds the final response, the query plus its list of sources, marks the job as done, and publishes the completion event, which wakes the backend instantly.

    The lesson

    The main lesson for us was simple. The most powerful tool is not always the right default. It is better to match the tool to the job you do most often, and to keep the heavy tool only for the few cases that really need it.

    Our first setup worked for a single search but broke when many searches ran at once, which is exactly what happens in a real product. The fix was not a bigger server or a clever trick. It was stepping back and asking what the job actually needed. Most of the time, the answer was less than what we first reached for.