Vacancy description
Hupo Zhongjian
India, Delhi
Full Stack Engineer Full Time Permanent in Delhi, India is listed on Jobeax. Browse 30,000+ vacancies available.
We are an AI-native start-up building sales enablement products in the banking, financial services, and insurance industry; Title: Senior Full Stack Engineer - AI Products
Employment Type: Full-time
We build AI products for the people who sell insurance and banking products across Asia. An advisor can practise a hard client conversation against an AI persona, get a nudge during a live call, or have an AI agent make the first call for them. Our clients are big, regulated, and spread across markets and languages - which means the interesting problems are latency, reliability, testing things that are not deterministic, security that satisfies a bank, and making an AI sound natural in Thai or Cantonese, not just English.
An advisor practises a pitch against an AI customer; a live call gets a real-time nudge; an AI agent phones someone and holds a conversation. Underneath all of that is a voice path - streaming audio in, speech-to-text, a language model, text-to-speech, audio out - running across languages including Thai, Bahasa Indonesia, Cantonese, Mandarin, Vietnamese and English.
Different products built it slightly differently, provider choices were made case by case, and when we add a language the quality check is whoever on the team happens to speak it. We want a senior engineer to own the voice path properly: turn-taking, interruptions, latency, which provider we use for which language and why, monitoring, and what happens when a provider has a bad day. And we want adding a language to be a process with test data and acceptance criteria, so provider and configuration changes can be evaluated consistently. Put streaming speech-to-text, language-model and text-to-speech providers behind a clean, configurable architecture.
Build the framework that tells us which provider is better for a given language: accuracy, latency, cost, and how it actually sounds.
Turn recordings and transcripts from native speakers into regression tests that run whenever a provider, prompt or configuration changes.
Get language, provider and tenant configuration under version control so it is reviewable and the same in every environment.
Run and write up incidents on the voice path, and turn what you find into permanent fixes.
Work with Product and QA to define what good voice quality means for a language and a use case.
Month one: own the regression set for one language and the map of what actually runs in each environment; Month two: provider benchmark harness working for at least two languages; a written procedure for onboarding a language, used once for real; 5 or more years building and running production backend or full-stack systems, with at least two on voice, audio, communications or something else where latency really matters.