We are looking for an experienced AI Engineer / Full-Stack Developer to build an AI-powered voice agent capable of having natural, real-time conversations with users.
The goal is to create a reliable voice-based AI assistant that can understand spoken language, respond naturally, maintain conversation context, and perform actions based on user requests.
The AI voice agent should be able to:
- Receive and process real-time voice input
- Convert speech to text (STT)
- Understand user intent using an LLM
- Generate natural, context-aware responses
- Convert ai responses back to speech (tts)
- support low-latency, real-time conversations
- maintain conversation history and context
- handle interruptions and follow-up questions naturally
- trigger external apis and business logic when necessary
- transfer conversations to a human operator when required
- store conversation logs and relevant metadata
- provide error handling and fallback behavior
potential technology stack
we are open to recommendations, but the project may involve technologies such as:
- backend:
node.js / Python
- ai/llm: openai api or equivalent llm
- frontend: react /
next.js
We are looking for someone with practical experience in:
- ai/llm application development
- voice ai or conversational ai
- openai or similar llm apis
- speech-to-text and text-to-speech apis
- real-time audio/websocket/webrtc systems
-
node.js or Python
- REST APIs and third-party integrations
- Database design
- Cloud deployment
Experience with Twilio, SIP, WebRTC, ElevenLabs, Whisper, function calling/tool use, RAG, or agent frameworks is a strong plus.
Please provide:
- Examples of AI voice agents or conversational AI systems you have built
- Your experience with real-time voice technologies
- Which STT/TTS/LLM stack you recommend and why
- Your preferred backend technology
- Estimated timeline
- Estimated budget
- Any architectural recommendations you would make for this project
Delivery term: Not specified