VVD Detect — the SDK
Who is speaking, answered for you. The detection engine proven in Virtual Video Director, offered to selected partners to build into broadcast switchers, conferencing systems, PTZ controllers and streaming appliances.
In any room with more than one microphone, the hard problem is not capturing audio — it is knowing who currently holds the floor, reliably enough to cut a camera on. VVD Detect listens across every microphone and answers that in real time. It also reads how the room feels, and it scales from a handful of microphones to hundreds.
What it tells your application
- The active speaker — one stable channel to put on screen, held sensibly through the natural pauses in speech rather than flickering on every breath.
- Who else is talking, and when two people talk at once — the cue for a wide shot.
- Silence, with a live meter and speaking-likelihood for every channel.
- The room’s mood — an optional read of the emotional tone (calm, tense, excited) that can drive shot width. It follows whoever is speaking on its own, from the audio you already send — nothing extra to wire up.
Built-in intelligence, no cloud
AI models are built into the library — one that tells speech from paper shuffles, chairs and air-conditioning, and one that reads the room’s emotional tone. Everything runs on the machine. Nothing is sent anywhere — no audio, no metadata, no phone-home — which is what lets it run on the firewalled and air-gapped networks broadcast facilities insist on.
Scales from one room to a whole rack
The AI runs across every channel in a single batched pass, so the same engine serves a small conference room or an installation with hundreds of microphones. For the largest channel counts it ships in GPU-accelerated builds — for NVIDIA (CUDA) and for Intel/AMD (DirectML) — drop-in with the standard CPU build: the same interface, you simply choose which file to ship.
What you get
A single native library with a simple, language-neutral C interface — active-speaker detection, AI voice-activity and room-mood sensing, with the models inside it. The standard CPU build is one self-contained file, no runtime to install. It comes with the same sample application written five times (C, C++, C#/.NET, Python and Rust), so you can start from whichever is closest to your codebase, plus a full integration guide.
How licensing works
Evaluate for free. The library runs unlicensed in a limited mode — the complete product, capped to a small number of channels — so you can build and test against your own audio before committing. Each licensed machine then gets a credential bound to that machine and verified entirely on the device, so it keeps working on an air-gapped network.
Let’s talk
If automatic, operator-quality speaker detection would strengthen your product, we would like to show you what it can do against your own audio.
Toby Mills — Noise Productions Ltd
toby@np.co.nz · +64 274 597427