On September 9, NVIDIA announced an expansion of its NVIDIA AI for Media developer toolset for broadcast and streaming ahead of IBC 2026. The update strengthens features for detecting AI-generated video and enhancing image quality, while folding localization processing into Holoscan for Media, its platform for live production. At IBC, held September 11-14 in Amsterdam, NVIDIA is pursuing two goals at once: making individual AI processes more sophisticated, and building the infrastructure to connect them across broadcast applications. Still, the software available today and the products partners are still integrating sit at different stages of readiness, a distinction that matters for anyone evaluating deployment.
Detecting and Enhancing Video with AI
The Synthetic Video Detector (SVD), which flags video that may have been AI-generated, was already announced at SIGGRAPH earlier this year. This update brings improved accuracy and deeper integration into systems used by news organizations. NVIDIA says the tool now reaches 99.3% accuracy for text-to-video content and 97.7% for image-to-video content.
Those figures don't mean the tool can spot fakes with equal reliability across every type of video. NVIDIA's IBC announcement doesn't detail the video sample population or the thresholds used to reach these numbers, and the company positions SVD as one input for editorial teams reviewing footage rather than a final verdict. Dalet, which builds newsroom systems, is integrating SVD scores and metadata into a cloud-based verification workflow where staff review the results on screen — a design that routes AI judgments into human review, not around it.
Video enhancement tools have also expanded. Video Super Resolution (VSR) upscales footage while reducing noise, blur, and compression artifacts. A new operating mode lets users prioritize either processing speed or image quality, and the tool now supports 10-bit video. VSR is available through the Video Effects SDK and as a NIM microservice for easier integration.
Separately, Video Frame Generation (VFG) creates new frames between existing ones to smooth motion. NVIDIA says it can double or quadruple frame rate, and Ross Video is integrating it into Rio Replay for sports production, where it currently supports 6x slow motion generation, with 8x interpolation still in development. It's worth noting that VFG produces synthesized intermediate frames — not moments actually captured by the camera.
TrueHDR, a tool for expanding dynamic range, converts SDR video to HDR output in real time, reaching brightness levels up to roughly 2,000 nits. It can be combined with VSR upscaling and VFG interpolation to improve existing footage within the same processing pipeline. NVIDIA also offers 3D Body Pose, which estimates joint positions and angles from single-camera footage. Vizrt is using this in virtual studios to drive lighting effects like reflections and shadows that respond to a performer's movement — an example of video-derived data being used for production purposes beyond image quality itself.
Connecting Live Broadcasts: MXL and Localization
NVIDIA is integrating the Media Exchange Layer (MXL) into Holoscan for Media. MXL provides a common mechanism for software-based broadcast functions to exchange video, audio, and data with each other. Holoscan itself is a reference design and toolset combining GPU processing, networking, and application deployment; MXL handles the connections between applications within that framework.
For example, if separate applications handle video analysis and processing, simply running them on the same compute infrastructure doesn't automatically connect their outputs — a mechanism for passing material between apps is still required. NVIDIA says MXL integration reduces the work developers need to do connecting each application individually, making it easier to combine capabilities from multiple vendors. This reads as infrastructure work aimed at running AI processing and traditional broadcast functions on shared hardware.
Localization is another use case built around coordinating multiple processes within a live broadcast. NVIDIA is integrating its Content Localization technology into Holoscan, providing reference workflows for handling subtitles, translated audio, and dubbing. This also covers lip-synced video and on-screen text localization. The idea is to let broadcasters select only the capabilities they need for a given program or distribution target, reducing the burden of maintaining separate infrastructure per language.
LipSync, which matches mouth movement to a different audio track, has improved handling of scenes where part of the face is obscured. Active Speaker Detection helps identify who is speaking in scenes with multiple people. NDI is also using these technologies, working on extending a single video stream into real-time translation and lip-synced dubbing. This reflects the reality that live broadcast requires handling audio and video together — generating a translated script alone isn't sufficient.
Training Sports AI on In-House Footage
Sports Intelligence Playbooks is a set of procedures for leagues and media companies to fine-tune NVIDIA's open models using their own sports footage and annotations. It covers data preparation, evaluation, and deployment, supporting the development of models that understand sport-specific rules, players, and game context. Unlike quality-enhancement tools like VSR, this is about tuning how a model interprets the content of video itself.
In early testing NVIDIA shared, accuracy on multiple-choice questions rose from roughly 53% to 94%, while free-response evaluation scores rose from roughly 5.7% to 66%. However, even though the test videos were previously unseen, the question formats resembled those used in training. Multiple-choice and free-response are separate evaluations, and these results can't be directly extended to unfamiliar sports or to the full range of questions that might arise during a live broadcast.
For rights holders, the value lies in being able to fine-tune models using their own footage and annotations. But confirming real-world benefit will likely require evaluation using a broadcaster's own game footage and the actual questions they want answered. Improvement on questions similar to training data and the ability to handle judgment calls needed in production should be measured as separate things.
Distinguishing What's Shipping from What's Still in Development
Within the same announcement, VSR, SVD integration with Dalet, and 8x interpolation sit at different stages — available, being integrated, and in development, respectively. Here's how NVIDIA's September 9 language breaks down by feature:
| Feature | Stage as of announcement | Distinction to keep in mind |
|---|---|---|
| VSR | Available via SDK and NIM | Component availability is separate from completion of a full broadcast system |
| SVD integration with Dalet | Being integrated into news verification workflow | Don't conflate SVD's own availability with the state of Dalet's integration work |
| 8x interpolation for sports replay | In development | Distinct from the already-available 6x slow motion generation |
This breakdown is based on the "available," "being integrated," and "in development" language used in NVIDIA's IBC announcement. It isn't a table certifying full product release dates or completed customer rollouts, and it doesn't indicate when 8x interpolation will be finished. Reading the announced features as all deployable at once risks confusing the availability of a component with its integration into a finished product.
The openness of MXL's connectivity comes with similar caveats. NVIDIA's developer documentation for Holoscan describes a design for combining technologies on top of compatible NVIDIA infrastructure. The fact that third-party applications and services can be connected more easily doesn't mean the same setup runs identically regardless of GPU. Freedom to choose broadcast applications and the underlying compute requirements to run them are two separate considerations that coexist.
What remains for broadcasters to verify on the ground is the latency and GPU count required when running subtitles, audio, and video correction simultaneously. This announcement doesn't provide measured figures or concrete cost-reduction numbers for the full pipeline, including localization. If the necessary image quality and latency can be achieved on existing infrastructure, broadcasters can add these capabilities to their current production environments while expanding the languages and distribution channels they serve.
