Grounded AI for astronomy — so the assistant reasons from your papers, not its memory.
Pipelines · Agents · Retrieval · Semantic Engineering of Context. A hands-on workshop: stand up a retrieval engine over your own preprints and notes, and let an AI write correct, domain-accurate analysis code instead of hallucinating your methods.
A parsec is a distance. Take your AI the distance — ground it in what your field actually knows.
Physicist → 10M QPS at Uber → AI & Applied ML → ready to help catalog the solar system. I speak both science and engineering fluently — my job today is to hand you the engineering, so you can keep doing the science.
Founder — Westover Labs · 2× BS Rochester, MS + PhD (ABD) UCF, MBA Quantic ·
Ex-Uber / HERE / Postmates · westover.dev · [email protected]
*QPS = queries per second
(Yes, it's a backronym — astronomers can't resist an acronym.)
A parsec is the astronomer's distance unit. This deck takes your AI the distance — four moves, and the heart of it is R: grounding the model in your own knowledge.
You're a scientist, not a software engineer — and you don't want to become one. So how do you get an AI to write correct code for your methods, without babysitting it or learning to code like a software engineer yourself?
The answer isn't "AI writes code." It's grounding — and that's what the rest of this workshop builds, hands-on.
Can your agent make fewer of these obvious mistakes?
You've probably tried it: ask an assistant to write your analysis, and it produces confident code that gets the domain subtly wrong — the wrong convention, an invented field name, a half-remembered method. Exactly the part you would catch, and it can't.
Without grounding, three failure modes dominate:
The model's knowledge is frozen at a training cutoff and blurred across everything. It has never read your preprint or your group's conventions.
Hand the model the exact, verified facts from your literature at question-time. Now it reasons from your methods, not its fuzzy memory.
Grounded, the AI writes analysis and simulation code that uses the real methods of your field. That is the force multiplier.
The deliverable is still your code and your science. Grounding is what makes the AI trustworthy enough to actually accelerate it.
github.com/fedorets/lsst-extendedness — a working Fink alert pipeline, built fast
because the AI was grounded in the Fink field docs the whole way, not guessing at them.
extendedness = 1 − classtar; SSO reassociation via ss_object_id changes; Julian dates;
magnitude-to-flux. Get any of these wrong and the science is wrong.
The AI retrieved the Fink field docs before writing — so it used the actual field names and conventions, not plausible-looking inventions. That is the whole difference.
What's an "AI assistant"? Three layers people conflate. Understanding the layers shows why the same grounding trick works in Claude Code, Codex, a laptop Ollama setup — anywhere, at the harness level.
The model. Weights frozen at training. Next-token prediction. No memory between calls, no tools, no UI.
Just a function: text in, text out.
The program. Manages the loop: context window, system prompt, UI (chat, TUI). Crucially: exposes hooks (e.g. UserPromptSubmit) and tool-calling wiring.
Examples: Claude Code, Codex, opencode, ChatGPT app.
The loop. Harness + tools + a task loop that runs the model autonomously: read files, run commands, call APIs, observe, repeat.
Not just chat — action + feedback.
Tonight's trick (LSST RAG hook): The astronomy-grounding happens at the harness layer — a UserPromptSubmit hook injects facts before the LLM ever sees the prompt. Swap harnesses, keep the same hook. That's proof the three layers are real and separable.
A modern coding assistant doesn't answer in one shot. It thinks, acts, sees the result, and thinks again until the task is done — reading files, running your tests, fetching docs. Understanding this loop is most of what you need to use it well.
Autocomplete predicts the next tokens from the prompt. One shot, no contact with the world.
An assistant takes actions, sees consequences, and course-corrects — it can run your test suite and fix what it broke, because it observed the failure.
"Act" is exactly where the loop can reach into your knowledge — fetch the right fact from your papers before it writes. That's the R section, and it's the heart of this talk.
"Tool-use" is more mechanical than it looks — and knowing the shape of it demystifies the whole thing. You describe a tool; the model asks to use it; your code runs it and hands back the result; the model reads that and continues.
It only proposes a call, as text. Your code executes it and returns the result. That boundary is where your data, your permissions, and your safety checks live — you stay in control.
Two standard doorways already exist for pointing the assistant at your data: MCP (next slide) and a plain CLI. Most of you will reach for the CLI — and that's a first-class choice.
The Model Context Protocol is an open standard that lets an AI assistant reach your tools and read your data through one uniform interface — the "USB-C for tools". Point it at a search-your-papers tool and grounding becomes something the assistant can do on its own.
A conversational path. The assistant decides when to fetch a fact and pulls it mid-answer, so grounding happens without you lifting a finger. Great when you're already working in the AI.
MCP is powerful but it's a server you run. For a lot of science work a plain command-line tool is simpler and enough — we'll put both side by side in the R section and let you choose.
People love to show swarms of agents orchestrating each other. For a working scientist that's mostly a distraction — you won't run a control room. What actually pays off is a single, always-on assistant wired into how you already work.
A real-time reviewer: as you write analysis code, an assistant — grounded in your papers — flags issues on the spot. "This uses osculating elements; your own method note says proper elements." It's a colleague reading over your shoulder, not a pipeline you operate.
Good engineering habits — review, tests, small steps — are what the AI amplifies. The value is the habit; the AI just makes it cheap to keep.
One grounded assistant, close to your work, beats a dozen agents you have to manage. Start there; you may never need more.
A language model knows only what was in its training data, frozen at a cutoff, blended into its weights. It cannot cite your preprint or your group's method note. RAG fixes that by fetching the relevant text at question-time and putting it in front of the model.
The model becomes an open-book reasoner: it reasons over the papers you handed it, not just what it happened to memorize. Open-book is how it writes your methods correctly.
Fresh: add a paper to the corpus — no retraining.
Cheap: no training run to add a document.
Attributable: the answer points back to a source you can check.
Yours: it's grounded in your literature, not the internet's average.
Retrieval finds the right text. Generation uses it faithfully. Most RAG failures are retrieval failures wearing a generation costume — so most of this section is about retrieval quality.
There are two honest paths to a grounded assistant. Neither is "the right" one — they fit different working styles, and you can mix them.
Wire retrieval into the assistant (via MCP or a hook). You ask a question or request code; the AI fetches the right facts and grounds itself automatically. Best when you live inside the AI and want it seamless.
A plain command you run: ground "my question" → it prints the retrieved facts. No AI in
the loop required. Reproducible, pipeable, versionable. Most scientists reach for this — and they're
right to. It's transparent and it fits how you already work.
The retrieval engine underneath is identical. The CLI is not the "lesser" path — it's often the smarter default: you can see exactly what it retrieved before anything reaches a model.
Here's the good news for a science project: you do not need our stack. Reach for an off-the-shelf, low-process tool. For most research, your source material is already markdown, PDFs, and preprints — and that is enough.
We run a process-heavy engine because we are refining the craft of building engines for many domains and a live ontology. You just need to use one, over your own documents. Different job, much simpler.
Corpus: a folder of PDFs / markdown — your papers, preprints, method notes.
Index: an off-the-shelf RAG tool (or ~90 lines of Python — you'll see it).
Ground: point your assistant, or the CLI, at it.
No schema. No migrations. No ontology. Add a PDF, re-index, done.
Start with markdown and PDFs. You can always graduate to something richer — but most science teams never need to.
This is the demo that matters. Ingest your preprints and PDFs, build a small RAG, and now the assistant reasons and writes over knowledge that sits outside its training data — your field, your methods, last month's paper.
The killer everyday use: write the code your own paper describes. Grounded in your method section, the AI implements the numerical computation the way you specified it — the right asteroid-family clustering in proper-element space, not a plausible-looking guess.
Ground the AI in a published paper and it can write code that reproduces that paper's result. Trust and reproducibility, demonstrated. That's the next slide.
This is the north star — the demo worth building toward. Point the engine at a published paper, ground the assistant in it, and have it write code that reproduces the paper's result — the same asteroid-family membership, the same survey-completeness curve, the same numbers.
Reproducibility is what scientists actually care about. A grounded AI that regenerates a known result from its paper isn't a party trick — it's evidence the grounding is real, and a genuinely useful tool for checking and extending others' work.
The horizon: the way Lean makes a proof deterministic and checkable, we want a deterministic, verifiable way to process survey data — pipelines you can trust by construction. Grounded RAG-code is an early step toward that determinism, not the finish line.
Full paper reproduction is the destination, not tonight's demo. What follows is the small, honest version of the same idea — grounding in ~90 lines, running live, that you can clone and verify yourself in the room.
We built you a tiny, self-contained companion repo. No database, no embeddings library, no API keys, no network — just Python's standard library and a JSON file of astronomy facts you can read and edit. Clone it and run it in the room.
Two panels appear: what the model sees without grounding (your bare question) and with it (your question + the exact retrieved facts). That injected block is the whole trick.
CLI path: demo.py shows the retrieval before/after — no AI needed. This is
"Way B" you can read and trust.
AI path: the same ~90-line script is a hook that grounds Claude Code and Codex on every prompt — one script, both tools, unchanged. This is "Way A".
Open astro_kg.json, add a fact, re-run. That's the entire editing workflow — no
schema, no migration. Swap in your facts and it grounds your domain.
Dense (vector) search catches meaning — "extendedness" ≈ "morphology score".
Sparse (keyword) catches exact tokens — a field name, an error code, ss_object_id.
Run both, fuse the lists. Pure-vector silently misses literal identifiers; pure-keyword misses paraphrase. Hybrid is the pragmatic default.
Too big → you retrieve noise and burn context. Too small → you sever the context a fact needs. Overlap, and chunk on structure (sections, equations) where you can.
An embedding is a list of numbers — a vector, often 384 to 3072 dimensions — produced by a model so that similar meaning → nearby vectors. The model learns to place text in a space where geometric closeness encodes semantic closeness.
Similarity is usually cosine similarity — the angle between vectors, ignoring length.
No single dimension means "asteroid-ness". Meaning is distributed across all dimensions. You don't read an embedding; you compare it.
Different embedding models = different spaces. You cannot mix vectors from two models in one index. Dimensions, domain (code vs prose vs science text), and max input length all matter — benchmark on your data, not a leaderboard.
Exact nearest-neighbour over millions of vectors is too slow. Approximate Nearest Neighbour (ANN) trades a sliver of recall for orders-of-magnitude speed. For a small paper corpus you won't even need it — but it's there when you grow.
Many databases can store vectors alongside your regular tables, so semantic search sits next to the data, filters, and backups you already trust — no separate system to run.
One query: filter by metadata, rank by meaning. That's semantic search.
Vector search is built for recall — cast a wide net cheaply. But the top of that net is noisy. A reranker re-scores the candidates for precision, so the best chunks land in the few slots your context budget allows.
Bi-encoder embeds question and chunk separately — you can pre-index, so it's fast but coarse.
Cross-encoder embeds them together in one pass — far more accurate, too slow to run over millions, perfect over 50.
That's the whole trick: recall cheaply, then rerank the shortlist.
Reranking is often the single highest-ROI upgrade to a mediocre RAG system — more than a fancier embedding model. Fix retrieval order before you touch anything else.
Vector RAG retrieves unstructured text. A knowledge graph retrieves structured facts — entities and the relations between them. The strongest systems use both (GraphRAG) — but this is a "graduate to it" tool, not where you start.
Ask "what method contrasts with osculating elements?" and the graph traverses to the answer — not a fuzzy paraphrase.
Multi-hop: "what depends on the thing we deprecated?" is a traversal, not a similarity search.
Exact: entities have identity; no near-duplicate drift.
Explainable: the path is the citation.
Markdown and PDFs get you most of the way. A knowledge graph is worth it once your facts have rich relationships you keep needing to traverse — add it when you feel that pain, not before.
The context window is the model's working memory — everything it can "see" this turn: instructions, history, retrieved chunks, tool results. It's finite (tokens), and it's the scarce resource the whole system budgets around.
You can't paste your whole library — or 60 PB of survey data — into the window. Retrieval is how you put only the relevant slice in front of the model. RAG is context-window management.
Deciding what earns a slot in the window: retrieve the right chunks, rerank, inject facts surgically, drop stale history. More tokens ≠ better — irrelevant context distracts the model ("lost in the middle").
Every "remember to…" you can move out of a bloated prompt and into a surgical, retrieved fact makes the window cleaner and the assistant sharper.
Grounding means the answer is anchored to retrieved evidence, not to the model's parametric memory. Done right, it's the biggest single lever against hallucination — and the reason you can trust the code it writes.
Constrain: "answer only from the context; if it's not there, say so."
Cite: require chunk IDs / source lines so claims are checkable.
Abstain: if retrieval is empty or low-confidence, refuse instead of guessing.
When retrieval returns stale, empty, or wrong chunks, a well-behaved model doesn't error — it quietly falls back to memory and confabulates fluently. The answer looks right and is wrong. This is the exact failure a scientist must be able to catch.
Freshness stamps on chunks, a confidence floor, "no-context → abstain", and evaluate on retrieval quality (was the right chunk even fetched?) not just the answer text. A silent failure you can't see is the dangerous one.
Long assistant sessions fill the window. When they do, the session is compacted — summarized to make room — and anything not written down first can be lost in the summary.
"Save what matters to durable memory, then compact." The session is ephemeral; your knowledge engine is not. Write decisions and findings down as you go, and retrieval brings them back on demand.
This is the same reason the knowledge engine pays off twice: it grounds today's answer and becomes the durable record the next session retrieves from. The library compounds.
| Layer | Simplest option | Purpose |
|---|---|---|
| AI assistant | Claude Code or Codex | The reason→act→observe loop, in your terminal |
| Your corpus | A folder of PDFs / markdown | Your papers, preprints, method notes — the source of truth |
| Retrieval | Off-the-shelf RAG (or ~90 lines) | Ground the assistant in your corpus, with sources |
| How to reach it | A CLI tool — or MCP | Way B (scriptable) or Way A (conversational) |
| Grow into → | Vector DB, reranker, knowledge graph | Only when the corpus and the pain grow — not on day one |
All open, all self-hostable, mostly your own files. Start with the top four rows. You can add the rest the day you actually need it.
Clone westoverlabs/parsec-rag-demo, run demo.py, then open
astro_kg.json and drop in a fact from your own work. You'll have grounded an assistant in your
domain before the coffee's cold.
Ground the AI in your literature, and it writes your science's code correctly — so you spend your attention on the science, not the syntax.
| Term | Plain-language definition |
|---|---|
| RAG | Retrieval-Augmented Generation — fetch relevant text at question-time and feed it to the model, so it answers "open-book" instead of from memory. |
| Retrieval | The "find the relevant chunks" step. Most RAG failures live here. |
| Grounding | Anchoring the answer to retrieved evidence rather than the model's memory. The main lever against hallucination. |
| Hallucination | A fluent, confident answer that isn't supported by any source. What grounding is meant to prevent. |
| Knowledge engine | Your corpus + retrieval, together — the thing that grounds the assistant in your own documents. |
| Chunking | Splitting documents into passages (~200–800 tokens) so they can be embedded and retrieved individually. |
| Chunk overlap | Repeating a little text between adjacent chunks so a fact isn't severed at a boundary. |
| Embedding | A vector (list of numbers) representing a piece of text so that similar meaning → nearby vectors. |
| Vector / vector space | The high-dimensional space embeddings live in; geometric closeness ≈ semantic closeness. |
| Cosine similarity | Similarity as the angle between two vectors (ignores length). The usual retrieval score. |
| Vector search | Finding the chunks whose embeddings are nearest to the question's embedding ("semantic search"). |
| ANN | Approximate Nearest Neighbour — trade a sliver of accuracy for huge speed when searching millions of vectors. |
| Hybrid search | Combining dense (vector) and sparse (keyword) retrieval, then fusing the ranked lists. |
| Reranking | Re-scoring retrieved candidates with a cross-encoder to push the truly-relevant chunks to the top. |
| Bi- / cross-encoder | Bi-encoder embeds question and doc separately (fast, coarse); cross-encoder scores them together (slow, precise) — used to rerank. |
| Top-k | The k nearest chunks you retrieve (k ≈ 5–50). Higher k = more recall, more context cost. |
| Anisotropy | Raw LM embeddings crowding into a narrow cone, so everything looks similar and discrimination suffers. |
| Term | Plain-language definition |
|---|---|
| Agent / Agentic AI | An AI in a loop with tools that can take actions — read files, run tests, fetch data — not just generate text. |
| Agent loop (ReAct) | Reason → act (use a tool) → observe the result → repeat, until the task is done. |
| Tool-use / function calling | The model emits a structured request to use a named tool; your code runs it and returns the result. |
| Tool schema | The typed description of a tool (name, inputs, purpose) the model uses to decide when and how to use it. |
| MCP | Model Context Protocol — an open standard that lets an assistant reach your tools and data over a uniform interface ("Way A"). |
| CLI tool | A plain command-line program you run to retrieve facts — scriptable, transparent, no AI required ("Way B"). |
| Hook | A small program the assistant runs on every prompt — e.g. to inject retrieved facts. Grounding becomes automatic and invisible. |
| Claude Code / Codex | Terminal-based AI coding assistants that run the agent loop and support the same grounding hook. |
| Knowledge Graph | A store of entities and the relations between them; structured memory you can traverse. |
| GraphRAG | Retrieval that traverses a knowledge graph (multi-hop, exact, explainable) alongside vector search. |
| Reviewer agent | An always-on assistant that flags issues as you write — the one orchestration pattern worth adopting solo. |
| Local model | An on-device model (e.g. Ollama on Apple Silicon) — private and quota-free; data never leaves your machine. |
| Term | Plain-language definition |
|---|---|
| Context window | The model's working memory this turn — instructions + history + retrieved chunks + tool results. Finite, measured in tokens. |
| Context engineering | Deciding what earns a slot in the window (the "SEC" in PARSEC). More tokens ≠ better; irrelevant context distracts the model. |
| Token | The unit models read and bill in — roughly ¾ of a word. Windows and costs are counted in tokens. |
| Lost in the middle | Models attend best to the start and end of the window; facts buried in the middle get overlooked. Why ordering + reranking matter. |
| System prompt | The standing instructions given to the model before the conversation. Guidance — not a security boundary. |
| Parametric memory | What the model "knows" baked into its weights from training — frozen at a cutoff, not your live data. |
| Fine-tuning | Further-training the weights on your data. Good for style/skills; RAG is usually better for facts. |
| Compaction | Summarizing a long session to free window space. Save durable notes before it happens. |
| Preprint / corpus | Your source material — the papers, preprints, and notes the knowledge engine retrieves from. |
| pgvector | A PostgreSQL extension that stores vectors next to your tables, so semantic search lives in the database you already have. |
| Evaluation (eval) | Measuring whether retrieval fetched the right chunk — not just whether the answer reads well. |
AI-development help for science teams — I speak both science and engineering fluently, and I'd rather hand you the engineering than sell you a platform.
github.com/westoverlabs/parsec-rag-demo — RAG over a knowledge base in ~90 lines.2× BS Rochester, MS + PhD (ABD) UCF, MBA Quantic · Ex-Uber / HERE / Postmates