Most coverage of AI training work — including most of our own content — focuses on text evaluation: rating written responses, comparing language model outputs, evaluating written explanations. But a significant and growing category of AI training work involves audio, video, and image modalities. Here's what exists, which platforms offer it, and what it pays.
Image Annotation
Image annotation — the longest-established form of AI training work — involves labeling elements in images so that computer vision models can learn to recognize them. The specific techniques include:
- Bounding box annotation — drawing rectangles around objects of interest (covered in our dedicated explainer)
- Semantic segmentation — labeling every pixel in an image with its category
- Image classification — categorizing entire images into predefined classes
- Quality evaluation — assessing whether AI-generated images meet quality standards
Pay for image annotation ranges from $12-25/hr for basic bounding box work to significantly higher for medical imaging annotation requiring domain expertise (one HireFeed report cited $25-60/hr generalist and $80-120/hr specialist for medical imaging specifically).
Audio Work
Audio-related AI training tasks are growing rapidly alongside voice AI development:
- Transcription — converting audio to text, including handling accents, background noise, and overlapping speakers
- Speech evaluation — rating AI-generated speech for naturalness, clarity, and appropriateness
- Voice data collection — recording yourself reading specific scripts to create training data (particularly valuable for regional dialects like Swiss German, as we describe in our German speaker piece)
- Audio quality rating — assessing whether AI-generated audio meets quality benchmarks
Video Annotation
Video annotation is more time-intensive than image work but increasingly in demand for autonomous vehicle training, sports analytics, and video AI applications:
- Action recognition labeling — identifying and labeling actions occurring in video clips
- Object tracking — following objects across frames as they move through a video
- Video quality evaluation — rating AI-generated video content
Which Platforms Offer Multimodal Work
CrowdGen and DataAnnotation.tech both include image and audio tasks alongside text evaluation. Toloka's pivot has specifically emphasized multimodal data collection. RemoExperts includes audio and transcription projects. For specialized medical imaging work, SuperAnnotate and Labelbox (both primarily enterprise tools) have their own contractor pools, though access is more selective.
Multimodal AI training work is expected to grow faster than text-only evaluation through 2027 as AI capabilities in vision, audio, and video mature — making it a good category to develop familiarity with now rather than after the demand peak.
Which Platforms Offer Multimodal Tasks
Multimodal AI training tasks are most consistently available on Mercor (image and video evaluation tracks), DataAnnotation.tech (image description, audio transcription correction, video annotation), Alignerr (visual and audio AI evaluation), and RemoExperts (audio and speech data tasks). Outlier AI periodically opens multimodal projects. The availability is project-dependent and tends to open in waves — worth setting alerts or checking platform dashboards weekly when primary text-based tasks are slower.
Pay Premium for Multimodal Work
Multimodal evaluation tasks typically pay a premium over equivalent text tasks — image and video assessment is more cognitively demanding and requires spatial reasoning skills that text-only annotators don't have. On platforms where both text and multimodal tasks are available, experienced contractors report 15-30% higher effective hourly rates on multimodal tracks. If you have a visual, creative, or media production background, prioritise applying to multimodal tracks specifically on every platform.
Ready to Start?
Apply directly or explore our top-ranked platforms.