Valentyn Budanov

AI Solutions Architect · Production Applied AI Systems · Orchestration & Evaluation

Builds and operates production conversational AI, and the evaluation systems that decide what ships. Co-Founder & CTO of WholeCall, a live voice AI platform serving 200+ business clients.

Colorado Springs, CO· Remote· GitHub· LinkedIn· hello@budanov.us
01

Production voice AI at scale.

Designed and built WholeCall's stage-based orchestration core and its evaluation platform, and led the ~7-engineer team around them.

200+ business clients · 260+ assistant configurations · 60+ concurrent live calls · 11 production languages
Self-service configuration cut client onboarding from ~10 minutes to ~90 seconds
Built the OpenAI Realtime and Deepgram Agent v2 voice integrations; declined LangChain/LangGraph because deterministic voice workflows needed custom stage-based orchestration
Three MCP servers (73 tools) behind refusal-gated tool access
Built the access control and audit logging used in the platform's ISO 27001 certification
02

Evaluation as an engineering discipline.

Deterministic, configuration-aware checks do most of the judging; a bounded LLM layer classifies only what code can't.

An LLM-simulated caller runs the real assistant through a 56-scenario library; every fix replays the baseline's exact scenario-and-seed tuples, scored as pass rates with 95% confidence intervals
The LLM judge is validated against 37 hand-labeled calls: floors of 80% agreement and 90% test-retest. After a model migration failed the floors, reports print NOT COMPUTED where judge numbers would be
Fixes are judged on the canary's own traffic, and each diagnosis must survive an adversarial pass that tries to refute it. Below a minimum of live calls, the pattern is not closed
A failure-attribution cascade localizes each production failure to one of seven places in the stack: prompt text, stage routing, function calls, stage config, base prompt, input, or the backend
Co-author, "Multi-Agent Virtual Game Environments," a university monograph on AI evaluation methodology and multi-agent systems (2025, ISBN 978-617-8181-56-7)
WholeCall's reporting view: an evaluation report with a verdict beside every call.
WholeCall's reporting view · demo data.
03

Ships code you can run.

Dictate is an open-source macOS dictation app built solo end to end: on-device Whisper via Core ML, 25 notarized releases in eight weeks, 387 tests, and a release script that refuses to publish anything unnotarized or built from a dirty tree.

Dictate a second after release, a green check on the panel, 34 words inserted into Mail.

See Dictate.

Screenshots, the privacy story, and the download.

Background

In 2021, my young son and I wired up our own J.A.R.V.I.S., Iron Man's talking computer. A button, speech recognition, GPT-3, a synthetic voice — and my son was holding a conversation with a machine.

Before WholeCall:
QA Lead → Acting CTO · regulated London fintech (SWIFT/SEPA payment flows)
QA Automation Lead · fintech payment products, Kyiv
Master's · Kyiv National Economics University

Write me: hello@budanov.us