A retrieval-augmented generation (RAG) system with citation-bound, grounding-verified output.
Single-stage dense retrieval runs over a heading-aware vector index. Generation is constrained
to the retrieved context and audited deterministically before anything reaches the client.
Request path
-
Edge controls. FastAPI on ASGI/Uvicorn. A per-IP sliding-window rate limiter
(10 req/min per IP) and a global budget cap (500 requests/day) bound abuse and spend.
-
Layer 1 input guard. A deliberately conservative pattern classifier for
prompt-injection phrasings runs before any retrieval or model call and short-circuits
to a fixed refusal, so recognisable attacks cost zero tokens and zero network calls.
-
Dense retrieval. The query is embedded with
all-MiniLM-L6-v2
(384-d, L2-normalised, ONNX Runtime on CPU) and matched by cosine similarity against a Pinecone serverless index,
top-k = 15.
-
Context assembly. Retrieved chunks are wrapped in
<chunk heading="…"> tags inside a <resume_context> block.
The system prompt declares both the context and the user turn to be untrusted data, not instructions.
-
Constrained generation. Claude Haiku 4.5 (Anthropic Messages API) must return a
JSON object
{answer, citations[]}, answering only from the supplied context.
-
Layer 2 grounding verification. Every citation is resolved against the chunks
actually retrieved for this request (exact match, or an unambiguous whole-segment breadcrumb suffix).
No valid citation, or unparseable or truncated output, and the model's text is discarded in favour of a
fixed refusal. Safety does not depend on the model choosing to behave.
Indexing path
The source document is split by a structure-aware chunker (markdown heading breadcrumbs, with long
bullet lists split into groups of ≤ 3 bullets), embedded as heading + body, and upserted to
Pinecone (cosine). A CI workflow re-indexes whenever the source document changes.
Design decisions, driven by measurement
-
Heading-augmented embeddings. Bullets rarely restate the employer they belong to, so
body-only embeddings made "what did you do at X?" unretrievable. Prepending the heading breadcrumb
before embedding fixed recall.
-
ONNX Runtime instead of PyTorch for embeddings. The same model weights held about
510 MB resident under PyTorch, over the 512 MB limit of a free-tier host. On ONNX Runtime they run in
about 210 MB and start in seconds rather than a minute. Vectors match the PyTorch ones to cosine
1.000000, but only after overriding the runtime's silent 128-token truncation back to the model's 256
(a setting the library ignores if passed normally). A test pins it.
-
Chunk granularity. Mean-pooled embeddings of long multi-topic chunks were nearly
orthogonal to focused queries (cosine ≈ 0.05 to 0.09). Splitting into small bullet groups restored ranking.
-
Recall over precision at k = 15. Dense similarity has no notion of recency, so
"employer before X" ranked the correct role 7th. At this corpus size, over-retrieving is cheaper
than missing it.
-
No similarity-threshold refusal gate. Cosine scores for legitimate questions varied
widely with phrasing (roughly 0.2 to 0.65), so a fixed cutoff would trade false refusals against
missed off-topic queries. Refusal is enforced structurally via citation verification instead.
-
No reranker or hybrid (BM25) stage. Deliberately omitted: at this corpus size
top-k = 15 already spans a large share of the index, so a second ranking stage was judged not worth
the added latency and complexity.
Guardrails and evaluation
-
Defense in depth: browser pre-check → server pre-check →
grounding verification. A shared must-flag / must-pass phrase corpus is run against both the Python
and JavaScript filters so they cannot drift apart.
-
Eval-driven development: pytest covers the deterministic layers with the model and
vector store stubbed to prove short-circuit paths make no network calls; promptfoo runs black-box evals
over HTTP against the live endpoint (grounding, refusal, injection, citation presence, voice);
node:test covers the browser guard. The suites run in CI, and repeated-run sampling
exposed two intermittent failures (output truncation, shortened citations) that single runs missed.
-
Cold-start mitigation on a scale-to-zero host: a
/health prefetch on page
load plus an ASGI lifespan hook that pre-loads the embedding model and index connection.
Stack
Python 3.13
FastAPI
Uvicorn
ONNX Runtime
Pinecone
Claude Haiku 4.5
Vanilla JS
GitHub Pages
Render
GitHub Actions
pytest
promptfoo
node:test