Move agent quality forward across the platform. Eval pipelines, guardrail design, multi-provider routing, and the model-side decisions that affect cost and latency for every tenant.
About the team
AI Research is a small, senior team building one of the most consequential layers of the Apragya platform. You'll work directly with the founders and the engineering / product leads - decisions are made in days, not quarters.
What you'll do
- Design and ship agent runtime improvements - HITL gates, guardrail evaluation, cost-aware orchestration
- Build evals, golden tests, and traces that catch agent quality regressions before customers do
- Partner with platform engineering on inference cost, latency, and provider routing
- Stay current on the model-side state of the art and bring back what matters
What we're looking for
- Hands-on experience building with LLMs in production (RAG, tool use, eval pipelines, agent loops)
- Strong engineering chops - Python, async patterns, and infra that handles real load
- Familiar with current research on agent quality, evaluation, and reliability
- Pragmatic about cost, latency, and the gap between a notebook demo and a shipping product
Nice to have
- Prior experience in early-stage startups - you've seen what works and what doesn't
- Open-source contributions or technical writing that shows how you think
- Familiarity with our stack (Next.js 14, FastAPI, PostgreSQL, Celery, Claude API)
What we offer
- Meaningful equity - everyone has skin in the game
- Comprehensive health insurance for you, your spouse, parents, and kids
- Remote-friendly with optional offices in Chennai and Bangalore
- Learning budget for courses, conferences, books, certifications
- Best-in-class gear and a parental leave policy that respects your life outside work