Revolutionary Real-Time Interview App Transforms Live Speech to Text with AI-Powered Analysis and Instant Feedback

September 28, 2026
Revolutionary Real-Time Interview App Transforms Live Speech to Text with AI-Powered Analysis and Instant Feedback
  • We’re building a real-time interview app where speech is converted to text live, captions update as the interviewer speaks, and agent replies stream as text fragments that are turned into natural-sounding speech sentence-by-sentence, with interruption handling and time-phase prompts to keep the flow smooth.

  • Skills are configured through JSON files—adding a new skill requires only dropping a JSON file with fields like Id, Title, InterviewInstruction, AgentTone, MaxPoints, InterviewDurationInMinutes, and PointInstructionMap, with no code changes needed.

  • The architecture relies on in-memory, vendor-agnostic components and exposes interfaces such as ISkillConfigRepository, ISessionInterviewService, IReportEvaluator, ISpeechConverter, IRealtimeInterviewAgent, IAudioChannel, and IRealtimeInterviewRunner.

  • The project envisions an app that conducts real-time voice interviews to assess a participant’s skill and delivers a comprehensive report card at the end.

  • Integrations use OpenAI for both the interviewer and the report agent (supporting streaming and structured JSON outputs) and Azure Speech for speech-to-text and text-to-speech, with efficiency enhancements like caching and per-voice synthesizers.

  • The system is split into two parts: Sessions/Skills (config loading, session lifecycle, in-memory storage) and Realtime (voice conversation, transcription, streaming agent responses, speech synthesis).

  • A WebSocket-based voice channel operates separately from the UI, with a Blazor-based browser component and a custom tag handling audio I/O and live captions.

  • Deployment and testing require a GitHub repo, OpenAI and Azure Speech keys, .NET 9, secrets management, and local run instructions.

  • Post-interview, a second AI agent analyzes the transcript to produce a score, label, summary, notable strengths, and areas for improvement.

  • The runner continuously listens, processes partial and final speech segments, supports interruption via cancellation tokens, and maintains per-interview state.

  • Interview flow is one question at a time with adaptive difficulty, allows interruptions or pauses, features live captions and transcripts, and runs with timed sessions that can grant extra time or end early.

  • Extensibility is achieved by dropping JSON files to add new skills and by configuring models, voices, and timing settings to suit different use cases.

Summary based on 1 source


Get a daily email with more Tech stories

More Stories