1 / 31
P A R S E C

Build your own
knowledge engine

Grounded AI for astronomy — so the assistant reasons from your papers, not its memory.

Pipelines · Agents · Retrieval · Semantic Engineering of Context. A hands-on workshop: stand up a retrieval engine over your own preprints and notes, and let an AI write correct, domain-accurate analysis code instead of hallucinating your methods.

A parsec is a distance. Take your AI the distance — ground it in what your field actually knows.

SpeakerJames Westover · Westover Labs
Date2026-08 (draft)
FormatTalk + hands-on repo
Westover Labs robot mascot
2 / 31
Who's Talking

James Westover

Physicist → 10M QPS at Uber → AI & Applied ML → ready to help catalog the solar system. I speak both science and engineering fluently — my job today is to hand you the engineering, so you can keep doing the science.

Petabyte-scale pipelines RAG & knowledge graphs Grounded AI assistants Python · scientific data Applied ML

Founder — Westover Labs · 2× BS Rochester, MS + PhD (ABD) UCF, MBA Quantic · Ex-Uber / HERE / Postmates · westover.dev · [email protected]
*QPS = queries per second

3 / 31
The Map

What PARSEC stands for

(Yes, it's a backronym — astronomers can't resist an acronym.)

A parsec is the astronomer's distance unit. This deck takes your AI the distance — four moves, and the heart of it is R: grounding the model in your own knowledge.

P
Pipelines
The starting point — a real LSST/Fink alert pipeline, and where AI actually helps a science team.
A
Agents
The assistant — how a coding AI reasons, uses tools, and reaches your data (via MCP or a CLI).
R
Retrieval — the heart
Ground the AI in your papers and preprints so it writes correct, domain-accurate code. RAG, embeddings, vector search, knowledge graphs.
SEC
Semantic Engineering of Context
The discipline — what earns a slot in the window, keeping answers grounded, and not-hallucinating.
4 / 31
PPipelines

Science software is getting harder to build, not easier

Data
Exponential growth. LSST: 20 TB/night, ~60 PB over 10 years, ~10M alerts a night.
Sprawl
Analysis spans many languages, frameworks, and institutions.
Teams
Small teams, big ambitions, limited software-engineering bandwidth.
Time
Every hour spent fighting code is an hour not spent on the science.

The question this deck answers

You're a scientist, not a software engineer — and you don't want to become one. So how do you get an AI to write correct code for your methods, without babysitting it or learning to code like a software engineer yourself?

The answer isn't "AI writes code." It's grounding — and that's what the rest of this workshop builds, hands-on.

5 / 31
PPipelines

Why generic AI code-gen isn't enough

Can your agent make fewer of these obvious mistakes?

You've probably tried it: ask an assistant to write your analysis, and it produces confident code that gets the domain subtly wrong — the wrong convention, an invented field name, a half-remembered method. Exactly the part you would catch, and it can't.

Without grounding, three failure modes dominate:

Duplication
Re-deriving or re-creating things that already exist — duplicate code, files, tickets, boilerplate. Grounding keeps the AI aware of what's already there.
Slop
Plausible-but-wrong output that sounds confident but gets subtle details wrong — exactly where it fails your science. Grounding gives it verified facts to reason from.
Drift
Generated state or configuration drifting from the source of truth over time. Grounding ties outputs back to live facts, keeping them synchronized.
The gap

It doesn't know your field

The model's knowledge is frozen at a training cutoff and blurred across everything. It has never read your preprint or your group's conventions.

The fix

Ground it in your papers

Hand the model the exact, verified facts from your literature at question-time. Now it reasons from your methods, not its fuzzy memory.

The payoff

Correct code, without being a coder

Grounded, the AI writes analysis and simulation code that uses the real methods of your field. That is the force multiplier.

The deliverable is still your code and your science. Grounding is what makes the AI trustworthy enough to actually accelerate it.

6 / 31
PPipelines · proof it works

A real LSST alert pipeline, built grounded

github.com/fedorets/lsst-extendedness — a working Fink alert pipeline, built fast because the AI was grounded in the Fink field docs the whole way, not guessing at them.

Fink
Live LSST/ZTF alert broker
5
Pluggable alert sources
Real
Tested, deployable — not a toy
2 days
Idea to production-ready

The domain-specific parts

extendedness = 1 − classtar; SSO reassociation via ss_object_id changes; Julian dates; magnitude-to-flux. Get any of these wrong and the science is wrong.

Why grounding mattered here

The AI retrieved the Fink field docs before writing — so it used the actual field names and conventions, not plausible-looking inventions. That is the whole difference.

7 / 32
AAgents

Chat program, demystified

What's an "AI assistant"? Three layers people conflate. Understanding the layers shows why the same grounding trick works in Claude Code, Codex, a laptop Ollama setup — anywhere, at the harness level.

LLM

The model. Weights frozen at training. Next-token prediction. No memory between calls, no tools, no UI.

Just a function: text in, text out.

Harness

The program. Manages the loop: context window, system prompt, UI (chat, TUI). Crucially: exposes hooks (e.g. UserPromptSubmit) and tool-calling wiring.

Examples: Claude Code, Codex, opencode, ChatGPT app.

Agent

The loop. Harness + tools + a task loop that runs the model autonomously: read files, run commands, call APIs, observe, repeat.

Not just chat — action + feedback.

Tonight's trick (LSST RAG hook): The astronomy-grounding happens at the harness layer — a UserPromptSubmit hook injects facts before the LLM ever sees the prompt. Swap harnesses, keep the same hook. That's proof the three layers are real and separable.

8 / 32
AAgents

The AI assistant is a loop, not autocomplete

A modern coding assistant doesn't answer in one shot. It thinks, acts, sees the result, and thinks again until the task is done — reading files, running your tests, fetching docs. Understanding this loop is most of what you need to use it well.

01
reason
Plan the next step given the goal + what's known so far.
02
act
Call a tool: read a file, run a test, fetch a doc, look something up.
03
observe
Read the tool result back in and adjust.
04
repeat / stop
Loop until the goal is met or a stop condition fires.

Autocomplete vs assistant

Autocomplete predicts the next tokens from the prompt. One shot, no contact with the world.

An assistant takes actions, sees consequences, and course-corrects — it can run your test suite and fix what it broke, because it observed the failure.

Where grounding plugs in

"Act" is exactly where the loop can reach into your knowledge — fetch the right fact from your papers before it writes. That's the R section, and it's the heart of this talk.

9 / 32
AAgents

How the assistant reaches a tool

"Tool-use" is more mechanical than it looks — and knowing the shape of it demystifies the whole thing. You describe a tool; the model asks to use it; your code runs it and hands back the result; the model reads that and continues.

# 1. you describe a tool { "name": "search_my_papers", "input": { "query": "text" } } # 2. the model asks to use it { "tool": "search_my_papers", "input": { "query": "proper elements" } } # 3. your code returns the facts { "result": "Proper elements are nearly invariant..." }

The model doesn't run anything

It only proposes a call, as text. Your code executes it and returns the result. That boundary is where your data, your permissions, and your safety checks live — you stay in control.

You don't have to build this by hand

Two standard doorways already exist for pointing the assistant at your data: MCP (next slide) and a plain CLI. Most of you will reach for the CLI — and that's a first-class choice.

10 / 32
AAgents

MCP: a standard doorway to your data

The Model Context Protocol is an open standard that lets an AI assistant reach your tools and read your data through one uniform interface — the "USB-C for tools". Point it at a search-your-papers tool and grounding becomes something the assistant can do on its own.

┌────────────┐ MCP (JSON-RPC) ┌──────────────────────┐ │ AI assistant│ ◄────────────────► │ Your MCP server │ │ (Claude / │ │ ├─ search_papers │ │ Codex) │ │ ├─ get_field_doc │ └────────────┘ │ └─ query_catalog │ └──────────────────────┘

This is "way #1": MCP + AI

A conversational path. The assistant decides when to fetch a fact and pulls it mid-answer, so grounding happens without you lifting a finger. Great when you're already working in the AI.

Honest scope

MCP is powerful but it's a server you run. For a lot of science work a plain command-line tool is simpler and enough — we'll put both side by side in the R section and let you choose.

11 / 32
AAgents · one that earns its keep

You don't need a fleet. You need one good habit.

People love to show swarms of agents orchestrating each other. For a working scientist that's mostly a distraction — you won't run a control room. What actually pays off is a single, always-on assistant wired into how you already work.

The one example worth keeping

A real-time reviewer: as you write analysis code, an assistant — grounded in your papers — flags issues on the spot. "This uses osculating elements; your own method note says proper elements." It's a colleague reading over your shoulder, not a pipeline you operate.

Frame it as a practice, not a platform

Good engineering habits — review, tests, small steps — are what the AI amplifies. The value is the habit; the AI just makes it cheap to keep.

One grounded assistant, close to your work, beats a dozen agents you have to manage. Start there; you may never need more.

12 / 32
RRetrieval · the heart

RAG: retrieval-augmented generation

A language model knows only what was in its training data, frozen at a cutoff, blended into its weights. It cannot cite your preprint or your group's method note. RAG fixes that by fetching the relevant text at question-time and putting it in front of the model.

question ──▶ retriever ──▶ top-k facts │ facts + question ──┴──▶ LLM ──▶ grounded answer / code

The model becomes an open-book reasoner: it reasons over the papers you handed it, not just what it happened to memorize. Open-book is how it writes your methods correctly.

Why it beats fine-tuning for facts

Fresh: add a paper to the corpus — no retraining.

Cheap: no training run to add a document.

Attributable: the answer points back to a source you can check.

Yours: it's grounded in your literature, not the internet's average.

The two halves

Retrieval finds the right text. Generation uses it faithfully. Most RAG failures are retrieval failures wearing a generation costume — so most of this section is about retrieval quality.

13 / 32
RRetrieval

Two ways to work with your data

There are two honest paths to a grounded assistant. Neither is "the right" one — they fit different working styles, and you can mix them.

Way A

MCP + AI — conversational

Wire retrieval into the assistant (via MCP or a hook). You ask a question or request code; the AI fetches the right facts and grounds itself automatically. Best when you live inside the AI and want it seamless.

Way B

A CLI tool — scriptable

A plain command you run: ground "my question" → it prints the retrieved facts. No AI in the loop required. Reproducible, pipeable, versionable. Most scientists reach for this — and they're right to. It's transparent and it fits how you already work.

The retrieval engine underneath is identical. The CLI is not the "lesser" path — it's often the smarter default: you can see exactly what it retrieved before anything reaches a model.

14 / 32
RRetrieval

Build your own knowledge engine

Here's the good news for a science project: you do not need our stack. Reach for an off-the-shelf, low-process tool. For most research, your source material is already markdown, PDFs, and preprints — and that is enough.

Why ours looks heavier

We run a process-heavy engine because we are refining the craft of building engines for many domains and a live ontology. You just need to use one, over your own documents. Different job, much simpler.

A science-project starter

Corpus: a folder of PDFs / markdown — your papers, preprints, method notes.

Index: an off-the-shelf RAG tool (or ~90 lines of Python — you'll see it).

Ground: point your assistant, or the CLI, at it.

No schema. No migrations. No ontology. Add a PDF, re-index, done.

Start with markdown and PDFs. You can always graduate to something richer — but most science teams never need to.

15 / 32
RRetrieval · the payoff

Ground the AI in your papers → correct code

This is the demo that matters. Ingest your preprints and PDFs, build a small RAG, and now the assistant reasons and writes over knowledge that sits outside its training data — your field, your methods, last month's paper.

01
ingest
Drop your papers / preprints into the corpus.
02
ground
Ask a question — the engine hands the model the real facts, with sources.
03
write code
Now "write the analysis" produces code using your methods, not invented ones.

Methodology-sync code

The killer everyday use: write the code your own paper describes. Grounded in your method section, the AI implements the numerical computation the way you specified it — the right asteroid-family clustering in proper-element space, not a plausible-looking guess.

And the showstopper →

Ground the AI in a published paper and it can write code that reproduces that paper's result. Trust and reproducibility, demonstrated. That's the next slide.

16 / 32
RRetrieval · the showstopper

Reproduce a published result from its paper

This is the north star — the demo worth building toward. Point the engine at a published paper, ground the assistant in it, and have it write code that reproduces the paper's result — the same asteroid-family membership, the same survey-completeness curve, the same numbers.

01
point at the paper
Ingest the PDF — method, parameters, data description.
02
ground & generate
The assistant writes the analysis from the paper's own method, not a guess.
03
run & compare
Execute it against the data and check: do the numbers match the paper?

Why this is the wow moment

Reproducibility is what scientists actually care about. A grounded AI that regenerates a known result from its paper isn't a party trick — it's evidence the grounding is real, and a genuinely useful tool for checking and extending others' work.

Toward verifiable pipelines

The horizon: the way Lean makes a proof deterministic and checkable, we want a deterministic, verifiable way to process survey data — pipelines you can trust by construction. Grounded RAG-code is an early step toward that determinism, not the finish line.

Tonight's concrete first step →

Full paper reproduction is the destination, not tonight's demo. What follows is the small, honest version of the same idea — grounding in ~90 lines, running live, that you can clone and verify yourself in the room.

17 / 32
RRetrieval · hands-on

Hands-on: RAG over a knowledge base in ~90 lines

We built you a tiny, self-contained companion repo. No database, no embeddings library, no API keys, no network — just Python's standard library and a JSON file of astronomy facts you can read and edit. Clone it and run it in the room.

# clone + run, no setup git clone https://github.com/westoverlabs/parsec-rag-demo.git cd parsec-rag-demo python3 demo.py "Why use proper elements for asteroid families?"

Two panels appear: what the model sees without grounding (your bare question) and with it (your question + the exact retrieved facts). That injected block is the whole trick.

Both ways, one engine

CLI path: demo.py shows the retrieval before/after — no AI needed. This is "Way B" you can read and trust.

AI path: the same ~90-line script is a hook that grounds Claude Code and Codex on every prompt — one script, both tools, unchanged. This is "Way A".

Make it yours

Open astro_kg.json, add a fact, re-run. That's the entire editing workflow — no schema, no migration. Swap in your facts and it grounds your domain.

18 / 32
RRetrieval

The retrieval pipeline: chunk → index → query

01
chunk
Split papers into passages (~200–800 tokens) with overlap so a fact isn't cut in half at a boundary.
02
embed
Turn each chunk into a vector (an embedding). Store vector + source metadata.
03
index
Build a nearest-neighbour index for fast lookup (or, for a small corpus, just scan — the demo does).
04
query
Embed the question, fetch the top-k nearest chunks, optionally rerank, hand them to the model.

Hybrid retrieval wins

Dense (vector) search catches meaning — "extendedness" ≈ "morphology score". Sparse (keyword) catches exact tokens — a field name, an error code, ss_object_id.

Run both, fuse the lists. Pure-vector silently misses literal identifiers; pure-keyword misses paraphrase. Hybrid is the pragmatic default.

Chunking is a real design choice

Too big → you retrieve noise and burn context. Too small → you sever the context a fact needs. Overlap, and chunk on structure (sections, equations) where you can.

19 / 32
RRetrieval

Embeddings: meaning as geometry

An embedding is a list of numbers — a vector, often 384 to 3072 dimensions — produced by a model so that similar meaning → nearby vectors. The model learns to place text in a space where geometric closeness encodes semantic closeness.

"minor planet" ──▶ [0.02, -0.19, 0.44, …] "asteroid" ──▶ [0.03, -0.17, 0.41, …] ← close "coffee machine" ──▶ [-0.5, 0.30, 0.02, …] ← far

Similarity is usually cosine similarity — the angle between vectors, ignoring length.

What the numbers are not

No single dimension means "asteroid-ness". Meaning is distributed across all dimensions. You don't read an embedding; you compare it.

Choosing a model matters

Different embedding models = different spaces. You cannot mix vectors from two models in one index. Dimensions, domain (code vs prose vs science text), and max input length all matter — benchmark on your data, not a leaderboard.

21 / 32
RRetrieval

Vector search at scale: your database can do this

Exact nearest-neighbour over millions of vectors is too slow. Approximate Nearest Neighbour (ANN) trades a sliver of recall for orders-of-magnitude speed. For a small paper corpus you won't even need it — but it's there when you grow.

ANN
"Find the nearest vectors, fast, approximately." The scaling trick behind semantic search.
top-k
You ask for the k closest chunks (k≈5–50). Bigger k = more recall, more context cost.
metadata
Filter and rank in one query — "nearest chunks, but only from this author / this year".

If you already have SQL, you're most of the way there

Many databases can store vectors alongside your regular tables, so semantic search sits next to the data, filters, and backups you already trust — no separate system to run.

SELECT id, body FROM chunks WHERE source = 'my_papers' ORDER BY embedding <=> :question LIMIT 8;

One query: filter by metadata, rank by meaning. That's semantic search.

22 / 32
RRetrieval

Reranking: precision after recall

Vector search is built for recall — cast a wide net cheaply. But the top of that net is noisy. A reranker re-scores the candidates for precision, so the best chunks land in the few slots your context budget allows.

01
retrieve top-50
Cheap, wide, approximate.
02
rerank to top-6
A cross-encoder reads (question, chunk) together and scores true relevance.
03
feed the model
Only the 6 best chunks reach the prompt. Higher signal, lower token cost.

Bi-encoder vs cross-encoder

Bi-encoder embeds question and chunk separately — you can pre-index, so it's fast but coarse.

Cross-encoder embeds them together in one pass — far more accurate, too slow to run over millions, perfect over 50.

That's the whole trick: recall cheaply, then rerank the shortlist.

Cheap wins here

Reranking is often the single highest-ROI upgrade to a mediocre RAG system — more than a fancier embedding model. Fix retrieval order before you touch anything else.

23 / 32
RRetrieval

Beyond documents: the knowledge graph

Vector RAG retrieves unstructured text. A knowledge graph retrieves structured facts — entities and the relations between them. The strongest systems use both (GraphRAG) — but this is a "graduate to it" tool, not where you start.

Entity: Method: Proper Elements Obs: "Nearly invariant over Myr timescales" Relation: used_for ──▶ asteroid-family clustering Relation: contrasts_with ──▶ osculating elements

Ask "what method contrasts with osculating elements?" and the graph traverses to the answer — not a fuzzy paraphrase.

Why graphs complement vectors

Multi-hop: "what depends on the thing we deprecated?" is a traversal, not a similarity search.

Exact: entities have identity; no near-duplicate drift.

Explainable: the path is the citation.

For a science project: don't over-build

Markdown and PDFs get you most of the way. A knowledge graph is worth it once your facts have rich relationships you keep needing to traverse — add it when you feel that pain, not before.

24 / 32
SECSemantic Engineering of Context

Context window & context engineering

The context window is the model's working memory — everything it can "see" this turn: instructions, history, retrieved chunks, tool results. It's finite (tokens), and it's the scarce resource the whole system budgets around.

Why RAG exists at all

You can't paste your whole library — or 60 PB of survey data — into the window. Retrieval is how you put only the relevant slice in front of the model. RAG is context-window management.

Context engineering (the "SEC")

Deciding what earns a slot in the window: retrieve the right chunks, rerank, inject facts surgically, drop stale history. More tokens ≠ better — irrelevant context distracts the model ("lost in the middle").

Every "remember to…" you can move out of a bloated prompt and into a surgical, retrieved fact makes the window cleaner and the assistant sharper.

25 / 32
SECSemantic Engineering of Context

Grounding — and how it degrades

Grounding means the answer is anchored to retrieved evidence, not to the model's parametric memory. Done right, it's the biggest single lever against hallucination — and the reason you can trust the code it writes.

Grounding techniques

Constrain: "answer only from the context; if it's not there, say so."

Cite: require chunk IDs / source lines so claims are checkable.

Abstain: if retrieval is empty or low-confidence, refuse instead of guessing.

Failure mode

Silent grounding degradation

When retrieval returns stale, empty, or wrong chunks, a well-behaved model doesn't error — it quietly falls back to memory and confabulates fluently. The answer looks right and is wrong. This is the exact failure a scientist must be able to catch.

Make it scream, not whisper

Freshness stamps on chunks, a confidence floor, "no-context → abstain", and evaluate on retrieval quality (was the right chunk even fetched?) not just the answer text. A silent failure you can't see is the dangerous one.

26 / 32
SECSemantic Engineering of Context

The window resets. Your knowledge does not.

Long assistant sessions fill the window. When they do, the session is compacted — summarized to make room — and anything not written down first can be lost in the summary.

01
window fills
Approaching the token limit; compaction is triggered.
02
save first
Write durable state to your corpus / notes — before the summary is taken.
03
compact & resume
Summary + re-retrieved facts carry the assistant across the boundary.

The discipline

"Save what matters to durable memory, then compact." The session is ephemeral; your knowledge engine is not. Write decisions and findings down as you go, and retrieval brings them back on demand.

Your corpus is the memory

This is the same reason the knowledge engine pays off twice: it grounds today's answer and becomes the durable record the next session retrieves from. The library compounds.

27 / 32
Putting It Together

The stack you can build this week

LayerSimplest optionPurpose
AI assistantClaude Code or CodexThe reason→act→observe loop, in your terminal
Your corpusA folder of PDFs / markdownYour papers, preprints, method notes — the source of truth
RetrievalOff-the-shelf RAG (or ~90 lines)Ground the assistant in your corpus, with sources
How to reach itA CLI tool — or MCPWay B (scriptable) or Way A (conversational)
Grow into →Vector DB, reranker, knowledge graphOnly when the corpus and the pain grow — not on day one

All open, all self-hostable, mostly your own files. Start with the top four rows. You can add the rest the day you actually need it.

28 / 32
For Science Teams

Recommendations

Ground first
Build a small RAG over your papers before anything else. A folder of PDFs + an off-the-shelf tool. This is the multiplier — everything downstream gets more correct.
Then code
Let the grounded assistant write your analysis. It now uses your field's real methods. You review and sign off — you don't have to be the coder.
Pick your path
CLI or MCP — whichever fits you. Most start with the CLI; it's transparent and reproducible. Add MCP if you want it conversational.
Stay small
Don't over-build. Markdown + PDFs beat an ontology you don't need. Add vectors, reranking, a graph only when you feel the pain.

Start today, in 10 minutes

Clone westoverlabs/parsec-rag-demo, run demo.py, then open astro_kg.json and drop in a fact from your own work. You'll have grounded an assistant in your domain before the coffee's cold.

The one-sentence version

Ground the AI in your literature, and it writes your science's code correctly — so you spend your attention on the science, not the syntax.

29 / 32
Glossary · 1 of 3 · Retrieval & RAG

Plain-language terms: retrieval

TermPlain-language definition
RAGRetrieval-Augmented Generation — fetch relevant text at question-time and feed it to the model, so it answers "open-book" instead of from memory.
RetrievalThe "find the relevant chunks" step. Most RAG failures live here.
GroundingAnchoring the answer to retrieved evidence rather than the model's memory. The main lever against hallucination.
HallucinationA fluent, confident answer that isn't supported by any source. What grounding is meant to prevent.
Knowledge engineYour corpus + retrieval, together — the thing that grounds the assistant in your own documents.
ChunkingSplitting documents into passages (~200–800 tokens) so they can be embedded and retrieved individually.
Chunk overlapRepeating a little text between adjacent chunks so a fact isn't severed at a boundary.
EmbeddingA vector (list of numbers) representing a piece of text so that similar meaning → nearby vectors.
Vector / vector spaceThe high-dimensional space embeddings live in; geometric closeness ≈ semantic closeness.
Cosine similaritySimilarity as the angle between two vectors (ignores length). The usual retrieval score.
Vector searchFinding the chunks whose embeddings are nearest to the question's embedding ("semantic search").
ANNApproximate Nearest Neighbour — trade a sliver of accuracy for huge speed when searching millions of vectors.
Hybrid searchCombining dense (vector) and sparse (keyword) retrieval, then fusing the ranked lists.
RerankingRe-scoring retrieved candidates with a cross-encoder to push the truly-relevant chunks to the top.
Bi- / cross-encoderBi-encoder embeds question and doc separately (fast, coarse); cross-encoder scores them together (slow, precise) — used to rerank.
Top-kThe k nearest chunks you retrieve (k ≈ 5–50). Higher k = more recall, more context cost.
AnisotropyRaw LM embeddings crowding into a narrow cone, so everything looks similar and discrimination suffers.
30 / 32
Glossary · 2 of 3 · Agents & Tools

Plain-language terms: the assistant

TermPlain-language definition
Agent / Agentic AIAn AI in a loop with tools that can take actions — read files, run tests, fetch data — not just generate text.
Agent loop (ReAct)Reason → act (use a tool) → observe the result → repeat, until the task is done.
Tool-use / function callingThe model emits a structured request to use a named tool; your code runs it and returns the result.
Tool schemaThe typed description of a tool (name, inputs, purpose) the model uses to decide when and how to use it.
MCPModel Context Protocol — an open standard that lets an assistant reach your tools and data over a uniform interface ("Way A").
CLI toolA plain command-line program you run to retrieve facts — scriptable, transparent, no AI required ("Way B").
HookA small program the assistant runs on every prompt — e.g. to inject retrieved facts. Grounding becomes automatic and invisible.
Claude Code / CodexTerminal-based AI coding assistants that run the agent loop and support the same grounding hook.
Knowledge GraphA store of entities and the relations between them; structured memory you can traverse.
GraphRAGRetrieval that traverses a knowledge graph (multi-hop, exact, explainable) alongside vector search.
Reviewer agentAn always-on assistant that flags issues as you write — the one orchestration pattern worth adopting solo.
Local modelAn on-device model (e.g. Ollama on Apple Silicon) — private and quota-free; data never leaves your machine.
31 / 32
Glossary · 3 of 3 · Context & building blocks

Plain-language terms: context

TermPlain-language definition
Context windowThe model's working memory this turn — instructions + history + retrieved chunks + tool results. Finite, measured in tokens.
Context engineeringDeciding what earns a slot in the window (the "SEC" in PARSEC). More tokens ≠ better; irrelevant context distracts the model.
TokenThe unit models read and bill in — roughly ¾ of a word. Windows and costs are counted in tokens.
Lost in the middleModels attend best to the start and end of the window; facts buried in the middle get overlooked. Why ordering + reranking matter.
System promptThe standing instructions given to the model before the conversation. Guidance — not a security boundary.
Parametric memoryWhat the model "knows" baked into its weights from training — frozen at a cutoff, not your live data.
Fine-tuningFurther-training the weights on your data. Good for style/skills; RAG is usually better for facts.
CompactionSummarizing a long session to free window space. Save durable notes before it happens.
Preprint / corpusYour source material — the papers, preprints, and notes the knowledge engine retrieves from.
pgvectorA PostgreSQL extension that stores vectors next to your tables, so semantic search lives in the database you already have.
Evaluation (eval)Measuring whether retrieval fetched the right chunk — not just whether the answer reads well.
32 / 32
Contact & Consulting

Westover Labs

AI-development help for science teams — I speak both science and engineering fluently, and I'd rather hand you the engineering than sell you a platform.

Hands-on
Clone the companion repo today: github.com/westoverlabs/parsec-rag-demo — RAG over a knowledge base in ~90 lines.
Workshops
Build-your-own-RAG workshops for research groups — ground your assistant in your own literature.
Retrieval
Knowledge-engine setup: index your papers, code, and data so the AI writes correct, domain-accurate code.
Advisory
Pick the right, smallest tool for your project — and skip the parts you don't need.
Writing
Writing the paper, not just the code? A grounded assistant can help draft it too — ask.

2× BS Rochester, MS + PhD (ABD) UCF, MBA Quantic · Ex-Uber / HERE / Postmates

Contents

← → arrows · space advance · T contents