A bug report is increasingly a screen recording rather than a carefully written sequence of steps. A site visit may be explained in a voice note while the details are fresh. A technician wearing smart glasses can point at the faulty part instead of stopping to describe what is already in view.
These are not poor versions of a written brief. They may be better evidence.
Yet many AI systems still ask somebody to turn the material into text first. Extract frames from the recording. Transcribe the audio. Write a prompt explaining what the camera already saw.
That extra translation creates work before the model has helped with anything.
The workforce has already changed
Gen Z is established in the workforce. Gen Alpha is following close behind. They grew up communicating through camera rolls, video calls, short clips and voice notes, not just documents and email.
Long before a child can read, they learn by seeing, listening, speaking, pointing and copying. Literacy is a remarkable cultural invention, but it is learned rather than an inborn human sense. Audio and video are closer to how people naturally take in and share the world.
That does not mean younger staff cannot write. It means writing need not be the default interface. Showing the problem while explaining it can be quicker and more precise than reconstructing the same event afterwards in formal prose.
The same matters for somebody with dyslexia, limited literacy, a language barrier or a disability that makes typing difficult. If the job requires them to convert clear visual or spoken evidence into a polished paragraph, the system is testing their writing before it addresses the actual problem.
Text remains useful. Many processes should still produce a searchable written record. It does not have to be the front door.
Some models can receive the original evidence
Google’s Gemini 3.7 Flash accepts text, images, video, audio and documents directly. Alibaba’s Qwen3.5-Omni can take those inputs together and respond in text or speech. Meta’s Muse Spark carries visual and audio information through tool-using work, while MiniMax M3 combines image and video input with coding and agent tasks.
Other model families still split the work between a main text model, transcription, extracted video frames and separate media services. That can work, but it is a pipeline somebody has to build and maintain. It can also discard the timing and connections between what was said, shown and heard.
Model selection therefore needs a basic question alongside price, speed and benchmark scores: what form does the information take when the job begins?
For a maintenance report, the useful source might be a recording of somebody pointing to an asset number, reading a gauge and capturing an unusual noise. A suitable system could prepare the ticket, attach the relevant clip, flag what it could not establish and ask the engineer to confirm the record.
Direct media input does not make the model’s conclusion automatically correct. It means the work can begin with the original evidence rather than a thinner conversion of it. A capable model is not the same as useful work, but a model that cannot receive the job’s real input begins at a disadvantage.
Design for the people arriving now
Smart glasses make the direction particularly clear. A person can show a problem from their own point of view while keeping both hands free. The spoken explanation, visible object and timing arrive together. Turning that into detached screenshots and a transcript before any useful work begins misses the point of the interface.
The broader shift is already visible outside work: the multimedia apocalypse already happened. Business systems will receive more screen recordings, voice notes, camera footage and wearable video, not less.
The sensible response is not to remove writing. It is to inspect how one recurring job really arrives, then choose a model and workflow that can meet staff there. Preserve the source, turn it into the record the business needs and keep a person at the consequential decision.
The workforce has caught up with multimedia. Many models and business systems are still catching up with the workforce.