# TTC-107 — Multimodal literacy: images, audio, and text

Treat every modality as partial evidence: describe what is observable, preserve source and accessibility alternatives, and separate observation from interpretation.

Level: beginner · Version: 1.0.0 · Last reviewed: 2026-10-03
Review status: reviewed

## Learning outcome
Compare image, audio, and text evidence while labeling observations, inferences, uncertainty, and accessibility needs.

## Explanation
Text can omit tone and layout; an image freezes one framed moment; audio lacks visual context and may be misheard. Models can add another layer of error through transcription, object recognition, localization, or inference. Record origin, capture conditions, transformations, and missing context. For images, distinguish pixels from inferred identity or intent. For audio, verify critical names, numbers, and speaker labels against the recording. Provide meaningful alt text, captions, transcripts, and text equivalents based on the user's purpose. Never infer sensitive traits merely from appearance or voice.

## Worked fictional example
Fictional case: a photo shows a wet floor beside a yellow sign; a voice note says ‘the west entrance is closed’; a text log says the sign was placed at 09:10. The assistant may report those observations, but it cannot conclude who caused the spill or that the entrance remains closed now.

## Reusable exercise
Use a synthetic scene represented by one image description, a short audio transcript, and a text record. Build a table of facts unique to each modality, agreements, conflicts, inaccessible details, and unresolved questions. Write an accessible combined account that does not erase uncertainty.

## Observable success criteria
- Observable details and interpretations are explicitly separated.
- Critical names, quantities, times, and speaker labels are checked against the original modality.
- The final account includes fit-for-purpose alt text, captions, or transcript and names missing context.

## Limitations
- Accessibility descriptions depend on purpose; one description will not serve every user or task.
- Synthetic or edited media may look authentic, and metadata can be missing or manipulated.

## Next review
Accessibility standards or common synthetic-media risks materially change.; A cited primary source is materially revised, replaced, or becomes unavailable.; Repeated learner results show that the exercise or success criteria are ambiguous.

## Copyable material
```text
# TTC-107 — Multimodal literacy
Objective: Use image, audio, and text as attributable partial evidence rather than complete reality.
Procedure: Record origin and transformations; label observation, inference, uncertainty, conflict, and missing modality.
Required evidence: Retain source references plus fit-for-purpose alt text, captions, transcript, and critical-detail checks.
Boundaries: Do not infer identity, intent, sensitive traits, or current state from ambiguous media.
Completion test: Another reviewer can trace each statement to a modality and identify what remains unknown.
Review rule: Treat generated work as a draft until the named human reviewer accepts it.
```

## Primary sources
- World Wide Web Consortium: Web Content Accessibility Guidelines (WCAG) 2.2 — https://www.w3.org/TR/WCAG22/
- World Wide Web Consortium: Making Audio and Video Media Accessible — https://www.w3.org/WAI/media/av/
- National Institute of Standards and Technology: Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile — https://www.nist.gov/publications/artificial-intelligence-risk-management-framework-generative-artificial-intelligence

Canonical URL: https://teachthecompany.com/school/multimodal-literacy/