Revolutionary Real-Time Interview App Transforms Live Speech to Text with AI-Powered Analysis and Instant Feedback
September 28, 2026
We’re building a real-time interview app where speech is converted to text live, captions update as the interviewer speaks, and agent replies stream as text fragments that are turned into natural-sounding speech sentence-by-sentence, with interruption handling and time-phase prompts to keep the flow smooth.
Skills are configured through JSON files—adding a new skill requires only dropping a JSON file with fields like Id, Title, InterviewInstruction, AgentTone, MaxPoints, InterviewDurationInMinutes, and PointInstructionMap, with no code changes needed.
The architecture relies on in-memory, vendor-agnostic components and exposes interfaces such as ISkillConfigRepository, ISessionInterviewService, IReportEvaluator, ISpeechConverter, IRealtimeInterviewAgent, IAudioChannel, and IRealtimeInterviewRunner.
The project envisions an app that conducts real-time voice interviews to assess a participant’s skill and delivers a comprehensive report card at the end.
Integrations use OpenAI for both the interviewer and the report agent (supporting streaming and structured JSON outputs) and Azure Speech for speech-to-text and text-to-speech, with efficiency enhancements like caching and per-voice synthesizers.
The system is split into two parts: Sessions/Skills (config loading, session lifecycle, in-memory storage) and Realtime (voice conversation, transcription, streaming agent responses, speech synthesis).
A WebSocket-based voice channel operates separately from the UI, with a Blazor-based browser component and a custom tag handling audio I/O and live captions.
Deployment and testing require a GitHub repo, OpenAI and Azure Speech keys, .NET 9, secrets management, and local run instructions.
Post-interview, a second AI agent analyzes the transcript to produce a score, label, summary, notable strengths, and areas for improvement.
The runner continuously listens, processes partial and final speech segments, supports interruption via cancellation tokens, and maintains per-interview state.
Interview flow is one question at a time with adaptive difficulty, allows interruptions or pauses, features live captions and transcripts, and runs with timed sessions that can grant extra time or end early.
Extensibility is achieved by dropping JSON files to add new skills and by configuring models, voices, and timing settings to suit different use cases.
Summary based on 1 source
