Surya Teja Reddy Dwarampudi

Member of Technical Staff 2 · Evaratus AI · Bengaluru, IN · open to conversations

Contact: hi@itssurya.com

Currently: building computer-use environments, automating agent evaluations, orchestrating agent workflows, building environment-generation tools.

Projects

01 · Automating agent evaluation (current)

Evaratus AI · agent evaluation · 2026–

Built an automated evaluation platform covering task generation, sandbox execution, failure diagnosis, and task repair. Owned the frontend, APIs, workflow orchestration, and cloud infrastructure.

Stack: Python, TypeScript, Temporal, Docker, AWS, Terraform

Role: Full-stack architecture · evaluation workflows · cloud infrastructure

The evaluation loop

I built a platform that turns application features into evaluation tasks, runs repeated agent trials in sandboxes, and collects the results for review.

When a trial fails

The workflow diagnoses failures and retries task repairs within a fixed limit. It flags environment issues for the team to resolve, then continues when an updated environment is ready.

The scope I owned

I built across the frontend, APIs, database, Temporal workflows, sandbox execution, and cloud infrastructure. Reusable snapshots and resumable runs preserve completed work through interruptions and deployments.

Resumable by design — Completed trials survive interrupted runs.

02 · An in-house CRM saving ₹5.4 Cr+ a year

Scaler · scale · 2024

Replaced LeadSquared with an internal CRM, cutting annual costs by 90%. Owned full-stack architecture and infrastructure, including OpenSearch across 3M+ leads and 100K+ monthly queries, role-based permissions, and a visual workflow builder.

Stack: Next.js, Turborepo, OpenSearch, PostgreSQL, React Flow, Kafka

Role: Full-stack architecture · infrastructure · lead search · workflows

Replacing a daily tool

The sales team used LeadSquared to manage leads and workflows. The internal replacement brought lead search, permissions, and configurable workflows into a single platform.

The scope I owned

I owned architecture across the frontend, backend, infrastructure, and deployment. The work included OpenSearch for large-scale lead data, permissions across 12+ modules, a React Flow workflow builder, and support for concurrent A/B tests.

Adoption and outcome

The team moved fully onto the CRM in four months. It indexed more than 3 million leads, served 100K+ monthly search queries, and reduced annual costs by 90% — ₹5.4 Cr+ in savings.

  • 90% — cost reduction
  • 3M+ — leads indexed
  • 100% — team adoption

03 · Computer-use environments for AI agents (current)

Evaratus AI · computer use · 2026–

Built computer-use environments for reinforcement learning and agent evaluation across multiple applications, with shared authentication, cross-app routing, and programmatic grading.

Stack: TypeScript, Python, Docker, MCP

Role: Environment architecture · agent orchestration · programmatic grading

Work across applications

Real tasks often span several tools. I built computer-use environments that bring applications together so an agent can complete a workflow across them.

Build the environment

I worked on container generation, build configuration, cross-app routing, shared authentication, MCP interfaces, and protected grading endpoints.

Verify the whole task

Programmatic graders inspect application state to check whether an agent achieved the goal. The example shows a calendar change followed by an update in a task tracker.

Cross-app task grading — Verify the final state across a complete workflow.

04 · Tools that build training environments (current)

Evaratus AI · developer tooling · 2026–

Built tooling for browser exploration, persistent interaction maps, coverage tracking, and agent-assisted environment creation with build and verification gates.

Stack: TypeScript, Python, Playwright, Docker

Role: Browser exploration · agent orchestration · coverage and verification

Explore the application

The tooling drives a browser to discover pages and interactions, recording them in a persistent state graph. Coverage tracking helps identify what still needs exploration.

Coordinate the build

I built orchestration for scoping work, planning routes, and coordinating frontend and backend generation, with checks before moving between stages.

Keep progress useful

Resumable exploration and deduplicated states let the workflow continue from what it has already learned. Agents can take a closer look at unclear interactions before the build proceeds.

Persistent exploration state — Resume with the application map and coverage intact.

05 · Real-time voice agent for sales calls

Scaler · flagship · 2025

Built a calling platform that handles sales conversations and books meetings. Combined Gemini, Deepgram speech recognition, and ElevenLabs speech synthesis, with Redis and WebRTC for call orchestration.

Stack: Gemini, Deepgram, ElevenLabs, Node.js, Redis, WebRTC

Role: Lead engineer · architecture, latency, orchestration

The task

Handle sales conversations and meeting booking through a real-time voice agent.

What I built

Led the architecture, latency work, and call orchestration. Combined Deepgram speech recognition, Gemini, and ElevenLabs speech synthesis with Redis and WebRTC.

The result

Average latency below 800 ms, more than 15 concurrent calls, and round-the-clock operation.

  • <800 ms — average latency
  • 15+ — concurrent calls
  • 24/7 — operation

06 · Call audit pipeline at 10K calls/month

Scaler · compliance · 2024

Built a pipeline that transcribes sales and support calls, checks them against compliance policies, and sends flagged passages to reviewers. Added review tooling and coaching feedback, with 94% precision on flagged issues.

Stack: Python, FastAPI, OpenAI, AWS MSK, Postgres, OpenSearch

Role: Pipeline architecture · LLM prompting · review tooling

The task

Check sales and support conversations against compliance policies and surface passages for human review.

What I built

Built the transcription and audit pipeline, LLM prompting, review tooling, and coaching feedback screens.

The result

Processed 10K+ calls per month with 94% precision on flagged issues and a three-day review-to-production loop.

  • 10K+ — calls/month
  • 94% — precision on flagged issues
  • 3 d — review-to-prod loop

07 · A practice ground for sales conversations

Scaler · training · 2025

Built a training platform where sales reps practice conversations with AI prospects. Created a library of 40+ personas, an evaluation harness, and feedback screens. The platform increased practice volume fivefold and cut ramp-up time by six weeks.

Stack: Gemini, TypeScript, Next.js, Postgres, pgvector

Role: End-to-end · prompts, eval harness, frontend

The task

Give sales reps a place to practice conversations with AI prospects.

What I built

Built the platform end to end: persona prompts, an evaluation harness, and the frontend for practice and feedback. Created a library of more than 40 personas.

The result

Practice volume increased fivefold and rep ramp-up time fell by six weeks.

  • 6 wks — rep ramp-up cut
  • 40+ — persona library
  • 5x — practice volume

About

A little background.

At Evaratus AI (formerly Scaler AI Labs), I work directly with frontier AI labs on the systems used to train and evaluate AI agents. I build computer-use environments for reinforcement learning and agent evaluation, including workflows that span multiple apps. My work also covers automated task generation, programmatic grading, and tools that create and validate new environments.

Previously at Scaler, I built user-facing products including its AI Mock Interview platform, alongside tools that helped sales and operations teams manage their work. I owned architecture across frontend, backend, infrastructure, and deployment.

I studied Computer Science at Bennett University, graduating in 2024 in the top 1% of my class with a 9.8/10 CGPA.

  • 3+ — frontier AI labs worked with directly
  • 100 K+ — daily users of products I built · Scaler

Experience

Member of Technical Staff 2 · Evaratus AI (formerly Scaler AI Labs) (current)

Feb 2026 to present · Bengaluru

  • Build computer-use environments spanning multiple applications for reinforcement learning and agent evaluation, working directly with 3+ frontier AI labs.
  • Own automated evaluation pipelines covering task generation, sandboxed agent execution, programmatic grading, failure diagnosis, and task repair.
  • Build agent-driven tooling for browser exploration, workflow discovery, and environment generation, with coverage tracking and verification.
  • Own full-stack engineering for evaluation platforms and internal analytics tools, including AWS infrastructure, data pipelines, and distributed evaluation runners.
  • Tested and delivered large-scale changes across 400+ RL environments through automated validation and coordinated rollouts.

Software Engineer I · Scaler by InterviewBit

Jul 2024 to Jan 2026 · Bengaluru

  • Built user-facing products including Scaler’s AI Mock Interview platform, alongside internal tools for sales and operations tracking and productivity.
  • Owned frontend, backend, and infrastructure architecture for an in-house CRM that replaced LeadSquared, cutting annual costs by 90% with full team adoption in four months.
  • Built OpenSearch lead search across 3M+ leads, role-based permissions across 12+ modules, and a visual sales workflow builder.
  • Developed real-time voice agents, LLM-powered call auditing, and AI simulations for sales training.

Software Engineer Intern · Scaler by InterviewBit

Mar 2023 to Jul 2024 · Bengaluru

  • Built mentor scheduling with Ruby on Rails and React, reducing no-shows from 11% to 3% and earning the Needle-Mover Award.
  • Built self-service admin tools so sales and operations teams could handle routine changes without engineering support.
  • Improved article delivery for 500K+ monthly active users with lazy loading and server-side pagination.

Full-Stack Developer Intern · SCSET · Bennett University

Sep 2022 to Jan 2023 · Greater Noida

  • Built Django dashboards for university staff to track and review inter-school activity.

Skills

  • AI agents & evaluation: Computer-use environments, Agent evaluation, Harbor, RL environments, MCP, Gemini / OpenAI / Anthropic, LangChain, Deepgram (ASR), ElevenLabs (TTS)
  • Languages: TypeScript, Python, JavaScript, Ruby, C++, SQL
  • Frameworks: Next.js, React, Node.js, FastAPI, Ruby on Rails, GraphQL
  • Libraries: Redux / RTK Query, React Flow, Prisma, Sentry, Temporal
  • Infrastructure: AWS, ECS / EC2 / RDS, OpenSearch, ClickHouse, MSK · Kafka, Docker, Terraform, GitHub Actions, Jenkins, Redis
  • Engineering practices: Agent orchestration, Event-driven systems, Role-based access control, A/B + feature flags, Workflow engines

Contact

Have something in mind? I’d love to hear about it.