About the Role
Own the intelligence layer - the AI research pipeline that classifies and risk-scores every tool we find. Discovery tells us what software an organization runs; the AI does the hard part: figuring out what each tool actually is, what it can do, how it handles data, and how risky it is. You'll own that research and enrichment system end to end - the LLM-backed agents, the prompts and context that drive them, the evals that keep them honest, and the cost and latency of running them at scale.
This is an engineering role, not a research one. You'll ship production TypeScript, and you'll be measured on the accuracy, cost, and reliability of the intelligence the product depends on.
What you'll work on
• The multi-agent researcher system: LLM-backed agents that research each tool across topics like platform, data policy, AI models, and agentic capabilities, and return structured, evidence-backed classifications.
• Evals and quality: design eval sets, measure classification accuracy and hallucination, and turn prompt changes into regression-tested, reviewable diffs instead of guesswork.
• Grounding and trust: cite evidence, resolve contradictions between AI output and validated data, and drive down hallucination on the fields that matter.
• Model routing and cost/latency: choose and route across providers, tune concurrency and caching, and keep the pipeline fast and affordable as volume grows.
• Structured outputs, tool/function calling, and the schemas and validation that make model output safe to persist.
• Deep observability into the pipeline — spans, traces, and metrics for every model call.
Tech you'll work with
TypeScript on Node 24 · LLM providers (Anthropic Claude, OpenAI, Gemini) · structured output & tool calling · prompt/context engineering · evals & regression harnesses · Sentry AI observability · Aurora Postgres + Drizzle for caching and results · AWS Lambda pipeline · an MCP server exposing our data to agents.
What we're looking for
• 3+ years of software engineering with hands-on, in-production LLM experience - you've shipped an AI-powered system that real users depend on, not just notebooks or demos.
• Strong prompt and context engineering: you treat prompts as artifacts you version, test, and improve.
• An eval-driven instinct: you reach for a measurement before you reach for a bigger model, and you know how to detect and reduce hallucination.
• Fluency with structured outputs, function/tool calling, and multi-agent orchestration.
• Solid engineering fundamentals — you build the pipeline around the model, not just call the API.
• Judgment about cost, latency, and provider trade-offs at scale.
Nice to have
• RAG / retrieval, embeddings, or agentic tool-building experience.
• Building eval and observability tooling for LLM systems.
• Security, compliance, or risk-scoring domain experience.
• Fluency with agentic coding tools as a force multiplier.