How the assistant works
← Talk to HeraThis is the write-up version of the chat widget on this site: what problem it's actually solving, what it's built with, and what separates it from a chatbot that just answers questions. Everything below describes what's shipped and running, not a roadmap.
The problem
A generic LLM chat box bolted onto a portfolio site is a parlor trick: it'll answer questions about Nelson with whatever it already knows (or, worse, confidently make something up) and it has no way to tell a visitor that it's guessing. Hera is scoped differently. It only answers from a curated corpus of this site's own content — the bio, the FAQ, and a write-up per project — retrieved fresh for every question rather than baked into a prompt once and left to go stale.
Ask it something the corpus doesn't cover and it says so instead of filling the gap. That's the actual product requirement: grounded answers with a source trail, not just answers.
What it's built with
Retrieval runs over pgvector in the same Postgres instance the rest of the site uses — no separate vector database to run or pay for. Embeddings are computed locally with fastembed (ONNX, BAAI/bge-small-en-v1.5, 384-dimensional), so indexing the content costs no API calls and no per-token fee.
Generation goes to Groq, and the agent's step-by-step reasoning and tool-calling loop is built on LangGraph: a bounded number of model↔tool round trips per turn, not one prompt-and-done call and not an unbounded loop that can run away on its own.
What makes it autonomous, not just conversational
A conversational bot answers the question you asked. An autonomous one decides what it needs to find out first. Hera already reaches into this site's own live projects when a question calls for it — pulling a real trade from the Trading Simulator, a real score from the Company Scorer, a real run from Pipeline World — instead of describing them from memory.
The clearest version of that is its flagship case: ask about a stock's risk, and instead of returning one canned reply, it chains its own sequence of tool calls — look up a live quote, feed that into a trend projection, then turn the projection into a written risk report — deciding at each step, based on what the last step actually returned, whether to keep going and what to call next. Nobody hard-coded that order. The model picks it, inside a bounded loop that stops it from spinning forever.
That loop is capped on purpose: a fixed maximum of tool round trips per turn, so "autonomous" means "decides its own path through a bounded set of steps," not "unsupervised and unbounded."
Curious what people actually ask it, and how often it answers from the corpus versus says it doesn't know? The stats page is built from the same request logs, with no model calls of its own. Or just ask Hera something directly.