
Speech-to-Action
Control your workflow with your voice.
Speech-to-Action is not dictation. Voice commands drive the diagnostic viewer and the reporter while your eyes stay on the images. It is the first step toward agentic radiology, built on a unified cloud-native platform.
How Sirona is Different
Speech is the natural clinical interface.
In radiology, the reading position is biomechanically constrained — your eyes must stay on the images, and your hands are occupied with input devices or measurement tools. Dictation solved the output problem: you can speak your report instead of typing. Speech-to-Action extends the voice interface beyond the report: navigate studies, control the viewport, and edit your report without leaving dictation. This is about clinical efficiency: focus on the images, reduce context switching, spend more time on clinical judgment.
What speech-to-action enables
Voice-driven control of the diagnostic viewer and the reporter today, with agentic voice in development — built on a unified data model and a unified event stream.
Real-Time Voice Commands
Navigate studies, control the viewport, and edit your report by voice, without leaving dictation. Commands respond in real time because viewer and reporter run on one platform.
Explore ReporterCortex Voice Agents (Coming Soon)
The next step we are building: invoking agentic AI through natural language, so a spoken instruction will trigger a chain of work rather than a single action.
Meet CortexUnified Event Stream
Workflow state changes — study assigned, claimed, dictated, signed — are written to one versioned event stream. That is the awareness layer agentic voice will be built on.
Event streamUnified Data Model
Viewer and reporter share one data model, so a voice command acts on the study that is actually open: the same study, series, and report in both.
Explore the data modelOne Platform, No Hand-offs
Viewer and reporter are one cloud-native platform, so a voice command is not a round trip between vendors' systems. Your command is your platform acting.
Cloud-nativePractice-Configurable Actions (Coming Soon)
Custom voice commands that trigger your own workflows are planned through the RadOS SDK, now in early access with design partners.
RadOS SDKThe sound of agentic radiology
Interface
Your voice is the fastest input device
Keyboard and mouse were designed for general-purpose computers. The reading position was never a general-purpose computer — it's a constrained, high-attention environment where your eyes belong on the images. Voice removes the friction. Voice commands drive the diagnostic viewer and the reporter, and they respond in real time.
Hands-free viewer navigation
Viewport control by voice
Report editing without leaving dictation
Stays in focus — eyes never leave the images
Context
Commands grounded in the session
Voice commands on Sirona act on the session in front of you: the study that is open, the series you are on, the report you are dictating. Because viewer and reporter share one data model, a command never has to ask which study you mean. Conditional clinical queries — 'if the priors mention liver nodules, hang the most recent one' — are what we are building next.
Acts on the study, series, and report that are open
Viewer and reporter state in one data model
No cross-vendor handoff per command
Grounded in the unified data model
Automation (Coming Soon)
From dictation to agentic voice
The direction is agentic: a voice command triggering a chain of work rather than a single action. Voice commands drive the reporter and the diagnostic viewer today; the agentic layer will chain them together.
One voice command will trigger a chain of actions
Cortex agents invoked by natural language
A practice-configurable command library via the RadOS SDK
Architecture
Why a unified platform comes first
On a multi-vendor stack, a voice command that crosses viewer, reporter, worklist, and AI is a coordination problem between vendors. Sirona owns the full stack, so the command stays inside one platform. That's the architectural prerequisite.
Unified data model + event stream
No cross-vendor API coordination
One system to act on — no vendor hand-offs
2
surfaces under voice control today: the diagnostic viewer and the reporter
Real-time
command response, because viewer and reporter are one platform
0
vendor boundaries for a voice command to cross
The platform behind agentic radiology
Why Agentic AI Requires RadOS
Watch as Dr. Mark Longo demonstrates the power of a new paradigm in radiology AI. Welcome to the age of agentic, embedded, multimodal, real-time, clinically aware AI assistants.
Dr. Mark Longo
Chief Technology Innovation Officer, Sirona
“Through Sirona's platform, Everlight will be able to build and deploy AI-powered automations across clinical, administrative, and operational workflows. This partnership will fundamentally transform how our radiologists practice, and position Everlight at the forefront of AI-enabled diagnostic medicine globally.”
Jeff Oakman
Global Chief Operating Officer, Everlight Radiology
FAQs
Isn't this just dictation?
Which commands are supported?
What microphone do I need?
Is this FDA-cleared?
How is this different from Dragon Medical?
Does this require Cortex?