Projects
Built an automated evaluation platform covering task generation, sandbox execution, failure diagnosis, and task repair. Owned the frontend, APIs, workflow orchestration, and cloud infrastructure.
Stack: Python, TypeScript, Temporal, Docker, AWS, Terraform
Role: Full-stack architecture · evaluation workflows · cloud infrastructure
The evaluation loop
I built a platform that turns application features into evaluation tasks, runs repeated agent trials in sandboxes, and collects the results for review.
When a trial fails
The workflow diagnoses failures and retries task repairs within a fixed limit. It flags environment issues for the team to resolve, then continues when an updated environment is ready.
The scope I owned
I built across the frontend, APIs, database, Temporal workflows, sandbox execution, and cloud infrastructure. Reusable snapshots and resumable runs preserve completed work through interruptions and deployments.
Resumable by design — Completed trials survive interrupted runs.
02 · An in-house CRM saving ₹5.4 Cr+ a year
Scaler · scale · 2024
Replaced LeadSquared with an internal CRM, cutting annual costs by 90%. Owned full-stack architecture and infrastructure, including OpenSearch across 3M+ leads and 100K+ monthly queries, role-based permissions, and a visual workflow builder.
Stack: Next.js, Turborepo, OpenSearch, PostgreSQL, React Flow, Kafka
Role: Full-stack architecture · infrastructure · lead search · workflows
Replacing a daily tool
The sales team used LeadSquared to manage leads and workflows. The internal replacement brought lead search, permissions, and configurable workflows into a single platform.
The scope I owned
I owned architecture across the frontend, backend, infrastructure, and deployment. The work included OpenSearch for large-scale lead data, permissions across 12+ modules, a React Flow workflow builder, and support for concurrent A/B tests.
Adoption and outcome
The team moved fully onto the CRM in four months. It indexed more than 3 million leads, served 100K+ monthly search queries, and reduced annual costs by 90% — ₹5.4 Cr+ in savings.
- 90% — cost reduction
- 3M+ — leads indexed
- 100% — team adoption
Built computer-use environments for reinforcement learning and agent evaluation across multiple applications, with shared authentication, cross-app routing, and programmatic grading.
Stack: TypeScript, Python, Docker, MCP
Role: Environment architecture · agent orchestration · programmatic grading
Work across applications
Real tasks often span several tools. I built computer-use environments that bring applications together so an agent can complete a workflow across them.
Build the environment
I worked on container generation, build configuration, cross-app routing, shared authentication, MCP interfaces, and protected grading endpoints.
Verify the whole task
Programmatic graders inspect application state to check whether an agent achieved the goal. The example shows a calendar change followed by an update in a task tracker.
Cross-app task grading — Verify the final state across a complete workflow.
Built tooling for browser exploration, persistent interaction maps, coverage tracking, and agent-assisted environment creation with build and verification gates.
Stack: TypeScript, Python, Playwright, Docker
Role: Browser exploration · agent orchestration · coverage and verification
Explore the application
The tooling drives a browser to discover pages and interactions, recording them in a persistent state graph. Coverage tracking helps identify what still needs exploration.
Coordinate the build
I built orchestration for scoping work, planning routes, and coordinating frontend and backend generation, with checks before moving between stages.
Keep progress useful
Resumable exploration and deduplicated states let the workflow continue from what it has already learned. Agents can take a closer look at unclear interactions before the build proceeds.
Persistent exploration state — Resume with the application map and coverage intact.
05 · Real-time voice agent for sales calls
Scaler · flagship · 2025
Built a calling platform that handles sales conversations and books meetings. Combined Gemini, Deepgram speech recognition, and ElevenLabs speech synthesis, with Redis and WebRTC for call orchestration.
Stack: Gemini, Deepgram, ElevenLabs, Node.js, Redis, WebRTC
Role: Lead engineer · architecture, latency, orchestration
The task
Handle sales conversations and meeting booking through a real-time voice agent.
What I built
Led the architecture, latency work, and call orchestration. Combined Deepgram speech recognition, Gemini, and ElevenLabs speech synthesis with Redis and WebRTC.
The result
Average latency below 800 ms, more than 15 concurrent calls, and round-the-clock operation.
- <800 ms — average latency
- 15+ — concurrent calls
- 24/7 — operation
Built a pipeline that transcribes sales and support calls, checks them against compliance policies, and sends flagged passages to reviewers. Added review tooling and coaching feedback, with 94% precision on flagged issues.
Stack: Python, FastAPI, OpenAI, AWS MSK, Postgres, OpenSearch
Role: Pipeline architecture · LLM prompting · review tooling
The task
Check sales and support conversations against compliance policies and surface passages for human review.
What I built
Built the transcription and audit pipeline, LLM prompting, review tooling, and coaching feedback screens.
The result
Processed 10K+ calls per month with 94% precision on flagged issues and a three-day review-to-production loop.
- 10K+ — calls/month
- 94% — precision on flagged issues
- 3 d — review-to-prod loop
07 · A practice ground for sales conversations
Scaler · training · 2025
Built a training platform where sales reps practice conversations with AI prospects. Created a library of 40+ personas, an evaluation harness, and feedback screens. The platform increased practice volume fivefold and cut ramp-up time by six weeks.
Stack: Gemini, TypeScript, Next.js, Postgres, pgvector
Role: End-to-end · prompts, eval harness, frontend
The task
Give sales reps a place to practice conversations with AI prospects.
What I built
Built the platform end to end: persona prompts, an evaluation harness, and the frontend for practice and feedback. Created a library of more than 40 personas.
The result
Practice volume increased fivefold and rep ramp-up time fell by six weeks.
- 6 wks — rep ramp-up cut
- 40+ — persona library
- 5x — practice volume