# Writing a Screen Reader in Rust

2024-11-20

I built Aria, a small Windows screen reader, to understand how an interface becomes speech.

## How I got here

Screen readers have always been fascinating to me.
I knew how to label a button, but I had only a vague idea of what happened between pressing Tab and hearing a screen reader announce it. I wanted to understand that part, so I built a small screen reader of my own.

## What it has to do

The first version had four jobs:

- Detect which control has keyboard focus.
- Read its name and type aloud.
- Respond to keyboard input, including a key to stop speech.
- Play short sounds for feedback.

I left out a graphical interface and navigation beyond the focused control. That kept the project small enough to build while learning the Windows APIs.

A full screen reader also needs ways to explore content that cannot receive focus, such as paragraphs and headings. Browsers expose this information through accessibility APIs; they don't simply hide it behind their sandbox. Supporting that navigation was beyond the scope of Aria. Chromium's [accessibility documentation](https://www.chromium.org/developers/accessibility/) is a useful starting point for how browsers expose their content.

## Reaching into the OS

Aria reads the information applications expose through Windows UI Automation. That includes properties such as a control's accessible name and type, along with events that report changes to the interface. The quality of those properties depends on the application providing them.

Rust was the systems language I was most comfortable with. The [uiautomation crate](https://github.com/leexgone/uiautomation-rs) gave me a wrapper around UI Automation, and the [windows crate](https://github.com/microsoft/windows-rs) provided the Windows Runtime APIs for speech synthesis and playback. I used `rodio` for the short feedback sounds and `mki` for keyboard input.

## Reading the screen

To follow focus, I registered a handler with UI Automation. The handler remembers the previous element so it can ignore consecutive events for the same control. Here is the setup, with the surrounding application code omitted:

```rust
struct FocusChangedEventHandler {
    previous_element: Mutex<Option<UIElement>>,
}

let automation = UIAutomation::new()?;
let handler = UIFocusChangedEventHandler::from(FocusChangedEventHandler {
    previous_element: Mutex::new(None),
});

automation.add_focus_changed_event_handler(None, &handler)?;
```

During testing, a single Tab press could produce several focus events for the same element. Announcing every event meant hearing the same button repeatedly. Comparing the element's runtime ID with the previous one let me skip those duplicates:

```rust
impl CustomFocusChangedEventHandler for FocusChangedEventHandler {
    fn handle(&self, sender: &UIElement) -> uiautomation::Result<()> {
        let mut previous = self.previous_element.lock().unwrap();

        if let Some(element) = previous.as_ref() {
            if element.get_runtime_id()? == sender.get_runtime_id()? {
                return Ok(());
            }
        }

        *previous = Some(sender.clone());
        drop(previous);

        let name = sender.get_name()?.trim().to_string();
        let control_type = sender.get_control_type()?;

        log::info!("Focus changed to: {name} ({control_type:?})");
        Ok(())
    }
}
```

<SameRuntimeIdDiagram />

This shortened handler stops at logging. In Aria, the next step is to assemble the announcement from the control's name and localized type. A button named Save becomes “Save, button.” The control type also determines whether to play a sound before speaking.

## Speaking

Windows provides the voice. Its [SpeechSynthesizer](https://learn.microsoft.com/en-us/uwp/api/windows.media.speechsynthesis.speechsynthesizer.synthesizetexttostreamasync?view=winrt-26100) turns a string into an audio stream, which a `MediaPlayer` can play. A minimal example using the Windows bindings looks like this:

```rust
let synthesizer = SpeechSynthesizer::new()?;

let stream = synthesizer
    .SynthesizeTextToStreamAsync(&HSTRING::from("Hello, World!"))?
    .get()?;

let source = MediaSource::CreateFromStream(
    &stream,
    &stream.ContentType()?,
)?;

let player = MediaPlayer::new()?;
player.SetSource(&source)?;
player.Play()?;
```

The `.get()` call waits for synthesis to finish. This shows the sequence, but it isn't enough for responsive speech: the application also needs to keep handling input while synthesis and playback are in progress.

## Earcons

An earcon is a short sound that communicates something about the interface. Aria plays one when focus reaches an editable field or combo box:

```rust
if control_type == ControlType::Edit || control_type == ControlType::ComboBox {
    play_sound(INPUT_FOCUSSED_SOUND);
}
```

I started with three sounds: startup, shutdown, and input focus. They are embedded in the executable with `include_bytes!`, so they don't need to be distributed as separate files:

```rust
pub const STARTUP_SOUND: &[u8] = include_bytes!("../assets/sounds/startup.mp3");
pub const SHUTDOWN_SOUND: &[u8] = include_bytes!("../assets/sounds/shutdown.mp3");
pub const INPUT_FOCUSSED_SOUND: &[u8] =
    include_bytes!("../assets/sounds/input-focussed.mp3");
```

The input sound gives a quick cue about the control before its name is read aloud.

## Listening for keys

The keyboard hook sends Escape directly to the speech controller. Other keys go to `on_keypress`, which handles spoken feedback while typing in an input. In the initial implementation, the binding looked like this:

```rust
mki::bind_any_key(Action::handle_kb(|key| {
    use Keyboard::*;

    match key {
        Escape => TTS::stop(true).unwrap(),
        _ => on_keypress(format!("{:?}", key)),
    }
}));
```

`TTS` is Aria's wrapper around synthesis and playback. Keeping those operations behind a shared interface lets both the keyboard and focus handlers control the same announcement.

## Putting it together

The difficult part was deciding what to do when input arrived during speech. If I moved from a text field to Save, the old announcement needed to stop and the new one needed to begin. Pressing Escape needed to stop speech without starting anything else.

<SpeechInterruptionDiagram />

That meant coordinating focus events, keyboard input, and audio playback. I moved speech work out of the event handlers and added a command-line interface and configuration file around it.

Aria was a learning project. Following focus covers only a small part of what people need from a screen reader, but it made missing labels much easier for me to notice. When a button has no useful name, there is very little to say about it.

The [source code is on GitHub](https://github.com/Coyenn/Aria) under the MIT license.
