aivarinnovations
Senior Voice AI Engineer Platform: Convogent Level: Mid–Senior (5+ years | 2+ years building voice AI in production) Location: Bangalore Office
You will own the distributed, real-time systems that orchestrate parallel voice calls at massive scale : the core platform behind every conversation Convogent runs. The job is keeping thousands of simultaneous calls fast, reliable, and natural, owning how STT/ASR, LLM, TTS, RAG, and tool-calling work together, and where every millisecond and every point of quality goes across the pipeline. What you’ll do Own voice AI at scale. Keep the pipeline fast and reliable at enterprise concurrency, thousands of simultaneous calls, by lifting calls-per-vCPU through extending Pipecat or rewriting hot paths in Go. Own end-to-end latency. Keep the conversation under ~800ms by decomposing the budget across STT, LLM, TTS, RAG, tool-calling, turn-detection, and network hops. Diagnose the bad call. When a call does not feel right, locate the fault in the pipeline fast by asking the right questions, reading the right signals, and owning the fix. Make quality measurable. Build voice-to-voice evals that gate releases and catch STT/LLM/TTS provider drift before customers do. Own provider tradeoffs. Decide how STT / LLM / TTS / RAG / tool-calling choices balance voice quality against latency, per conversation flow. Build adjacent products. Extend the platform into live assist, agent assist, and what comes next, on the same pipeline foundations. Set the bar. Review, mentor, and define engineering standards for a growing team. What we’re looking for 5+ years building production products and systems , with at least 2 of those years in real-time voice, operating it in production, not just shipping a demo. Deep understanding of the full voice pipeline , every stage end to end and how each one affects the next and shapes quality and latency. Strong in Go/Python , comfortable across the whole pipeline. Production-grade, performance-critical Go is especially valued: you have rewritten hot-path components in Go to lift throughput under real load. A track record of running voice AI at scale , scaling Pipecat (or a comparable real-time voice framework) to high concurrency in production. Provider-tradeoff fluency: how STT / LLM / TTS / RAG / tool-calling choices play off each other on quality and latency. An eval-driven approach to quality. You measure voice quality with real signals, not vibes, and have built or run evals on a real-time voice or LLM system. Fluency with agentic coding tools and a habit of treating quality as a measured signal. What we’re NOT looking for A prompt engineer / "AI app" builder who has only called LLM APIs and never operated a real-time pipeline under load. Someone who has only used managed platforms (Vapi/Retell/Bland) and never scaled the layer beneath them. A pure web-backend CRUD engineer with no real-time/streaming/voice exposure. An ML researcher who wants to train/fine-tune models. We integrate providers; we don't build models.