How the chat agent finds sources
The codex/agent-improvements branch gives the main chat agent one new tool, rebuilds the other two, and adds limits on how much research a single answer can do. This page explains what each tool does and how a turn runs.
A turn at a glance
Resolve names
lookup_catalog turns "Ibn al-Qayyim" or "Sahih al-Bukhari" into catalog IDs.Search
search runs a hybrid query, optionally filtered by those IDs or by an author death-year range.Read around
expand pulls the passages right before and after a promising result.AnswerWhen the research budget runs out, or step 18 is reached, the model must write its answer from what it has.
Steps 1–3 can repeat and run in any order. Every passage already retrieved in the turn is excluded from later searches, so each call can only bring back new material.
The three tools
lookup_catalog New
Resolves an author, book or genre name to canonical catalog IDs.
- Input:
kind(author|book|genre) and a free-textquery. - How it works: it runs a fuzzy match over the server's in-memory library index and returns up to 5 candidates, labelled in the user's language.
- What comes back: the ID and label, plus the author's death year (AH) for authors and books and the author and genre for books. Each candidate also gets an
indexedflag. To set it, the server checks the search index for at least one passage from that source, because the catalog's own flag can lag behind imports. - Rules for the model: the results identify sources but are not evidence and can't be cited. If a name is ambiguous, the model should use the question to pick, or ask the user.
- It does not count against the research budget.
lookup_catalog({ kind: "author", query: "ابن القيم" })
→ [{ kind: "author", id: 1234, label: "Ibn Qayyim al-Jawziyya",
deathYearAH: 751, deathInexactLabel: null, indexed: true }, …]
Illustrative output; the ID is made up.
search Rebuilt
Searches the corpus and returns up to 20 passages.
- Input:
query(Arabic or English), a UI-onlylabel, and optionalfilters. The oldmodeparameter (semantic vs keyword) is gone. - Always hybrid: each call runs a vector search and a BM25 keyword search side by side (24 candidates each) and merges them with reciprocal rank fusion. The Cohere reranker (
rerank-v4.0-pro) then picks the best 20. Passages are delivered whole, with no byte caps. - Filters the model can set:
authorIds,bookIds,categoryIds: up to 20 IDs each, normally taken fromlookup_catalog. An unknown ID is rejected with "resolve it with lookup_catalog".deathYearAH: a range on the author's death year (gt/gte/lt/lte), so the model can search "scholars before 300 AH". Authors with no recorded date are excluded, not matched by accident.
- How filters combine: the model's filters all apply together (AND). The user's own scope from the composer (selected authors, books or genres) still applies on top. A search can narrow that scope but never widen it. Filters restrict who wrote a passage, not who it mentions.
- No repeats: passages already retrieved earlier in the same turn are excluded at the database level. Repeating the same query and filters returns the cached result and counts as "no progress".
search({
query: "الصبر عند المصيبة",
label: "Patience in hardship",
filters: { authorIds: [1234], deathYearAH: { lte: 800 } }
})
expand Tightened
Reads the context around a passage the model already has.
- Input:
book_id,version_id,sequence_number, copied from a retrieved passage. Passages now carrysequence_numberso the model can do this. - Returns: up to 2 neighbouring passages on each side, down from 5.
- Guarded: coordinates the model guesses, rather than copies from a passage retrieved this turn or in history, are refused. The neighbours also inherit the filters of the search that found the original passage, so expanding can't leak outside the requested scope.
Research limits per turn
- 8 retrieval calls at most (
search+expandcombined). The slot is reserved before any I/O, so parallel calls can't overspend. - Stop when it stops helping: two calls in a row that return no new passages end the research.
- 20 agent steps overall. From step 18, tools are switched off (
toolChoice: "none") so the last steps are reserved for writing the answer. - Explicit stop signal: once a limit is hit, a tool call returns "Retrieval budget exhausted … Synthesize your answer from the evidence already retrieved" instead of an empty list. The model can't mistake the stop for "no sources exist". These are expected stops, so they're kept out of Sentry.
- Leaner context: a passage that comes back twice is sent to the model only once. Chat history is capped at the last 20 stored messages, and older reasoning is pruned.
Before vs after
| main | agent-improvements | |
|---|---|---|
| Tools | search, expand | + lookup_catalog |
| Search mode | Model picks semantic or keyword | Always hybrid (vector + BM25, fused) |
| Candidates → delivered | 50 semantic → rerank → 20 | 24 + 24 → fuse → rerank → 20 |
| Model-set filters | None (only the user's composer scope) | Author / book / genre IDs and death-year range, within the user's scope |
| Expand window | 5 each side, any coordinates | 2 each side, only from retrieved passages |
| Duplicates | Could resurface across calls | Excluded at query time; identical calls memoised |
| Budget | Step limit only | 8 calls, 2-strike no-progress stop, synthesis from step 18 |
Things to know before shipping
Prompt not promoted. The wording that teaches the model these tools is in
platform/main-gemini v28 (retrieval-v2). Production still serves v24. Local dev picks up v28 through the latest label. In the admin playground you have to select v28 explicitly; v24 is pinned to the top of the list.Death-year data. Date-range filters only work once the indexed passages carry
author_death_year. The v28 commit note says to promote the prompt only after the manual date backfill, which (per the branch's plan.md) hasn't been run yet.Settings are a starting point. The 24/20 numbers were not tuned on this branch. The earlier ablation tested 25 per branch, so quality parity for these exact settings hasn't been measured.