Ask me about my background

This is an AI assistant answering on Aby George's behalf, speaking as him. My answers come from my resume, with sources cited.

Technical architecture: how this agent works

A retrieval-augmented generation (RAG) system with citation-bound, grounding-verified output. Single-stage dense retrieval runs over a heading-aware vector index. Generation is constrained to the retrieved context and audited deterministically before anything reaches the client.

Request path

  1. Edge controls. FastAPI on ASGI/Uvicorn. A per-IP sliding-window rate limiter (10 req/min per IP) and a global budget cap (500 requests/day) bound abuse and spend.
  2. Layer 1 input guard. A deliberately conservative pattern classifier for prompt-injection phrasings runs before any retrieval or model call and short-circuits to a fixed refusal, so recognisable attacks cost zero tokens and zero network calls.
  3. Dense retrieval. The query is embedded with all-MiniLM-L6-v2 (384-d, L2-normalised, ONNX Runtime on CPU) and matched by cosine similarity against a Pinecone serverless index, top-k = 15.
  4. Context assembly. Retrieved chunks are wrapped in <chunk heading="…"> tags inside a <resume_context> block. The system prompt declares both the context and the user turn to be untrusted data, not instructions.
  5. Constrained generation. Claude Haiku 4.5 (Anthropic Messages API) must return a JSON object {answer, citations[]}, answering only from the supplied context.
  6. Layer 2 grounding verification. Every citation is resolved against the chunks actually retrieved for this request (exact match, or an unambiguous whole-segment breadcrumb suffix). No valid citation, or unparseable or truncated output, and the model's text is discarded in favour of a fixed refusal. Safety does not depend on the model choosing to behave.

Indexing path

The source document is split by a structure-aware chunker (markdown heading breadcrumbs, with long bullet lists split into groups of ≤ 3 bullets), embedded as heading + body, and upserted to Pinecone (cosine). A CI workflow re-indexes whenever the source document changes.

Design decisions, driven by measurement

  • Heading-augmented embeddings. Bullets rarely restate the employer they belong to, so body-only embeddings made "what did you do at X?" unretrievable. Prepending the heading breadcrumb before embedding fixed recall.
  • ONNX Runtime instead of PyTorch for embeddings. The same model weights held about 510 MB resident under PyTorch, over the 512 MB limit of a free-tier host. On ONNX Runtime they run in about 210 MB and start in seconds rather than a minute. Vectors match the PyTorch ones to cosine 1.000000, but only after overriding the runtime's silent 128-token truncation back to the model's 256 (a setting the library ignores if passed normally). A test pins it.
  • Chunk granularity. Mean-pooled embeddings of long multi-topic chunks were nearly orthogonal to focused queries (cosine ≈ 0.05 to 0.09). Splitting into small bullet groups restored ranking.
  • Recall over precision at k = 15. Dense similarity has no notion of recency, so "employer before X" ranked the correct role 7th. At this corpus size, over-retrieving is cheaper than missing it.
  • No similarity-threshold refusal gate. Cosine scores for legitimate questions varied widely with phrasing (roughly 0.2 to 0.65), so a fixed cutoff would trade false refusals against missed off-topic queries. Refusal is enforced structurally via citation verification instead.
  • No reranker or hybrid (BM25) stage. Deliberately omitted: at this corpus size top-k = 15 already spans a large share of the index, so a second ranking stage was judged not worth the added latency and complexity.

Guardrails and evaluation

  • Defense in depth: browser pre-check → server pre-check → grounding verification. A shared must-flag / must-pass phrase corpus is run against both the Python and JavaScript filters so they cannot drift apart.
  • Eval-driven development: pytest covers the deterministic layers with the model and vector store stubbed to prove short-circuit paths make no network calls; promptfoo runs black-box evals over HTTP against the live endpoint (grounding, refusal, injection, citation presence, voice); node:test covers the browser guard. The suites run in CI, and repeated-run sampling exposed two intermittent failures (output truncation, shortened citations) that single runs missed.
  • Cold-start mitigation on a scale-to-zero host: a /health prefetch on page load plus an ASGI lifespan hook that pre-loads the embedding model and index connection.

Stack

Python 3.13 FastAPI Uvicorn ONNX Runtime Pinecone Claude Haiku 4.5 Vanilla JS GitHub Pages Render GitHub Actions pytest promptfoo node:test