By Jonas Müller · · EarnWithAI.tech

Most coverage of AI training work — including most of our own content — focuses on text evaluation: rating written responses, comparing language model outputs, evaluating written explanations. But a significant and growing category of AI training work involves audio, video, and image modalities. Here's what exists, which platforms offer it, and what it pays.

Image Annotation

Image annotation — the longest-established form of AI training work — involves labeling elements in images so that computer vision models can learn to recognize them. The specific techniques include:

Pay for image annotation ranges from $12-25/hr for basic bounding box work to significantly higher for medical imaging annotation requiring domain expertise (one HireFeed report cited $25-60/hr generalist and $80-120/hr specialist for medical imaging specifically).

Audio Work

Audio-related AI training tasks are growing rapidly alongside voice AI development:

Video Annotation

Video annotation is more time-intensive than image work but increasingly in demand for autonomous vehicle training, sports analytics, and video AI applications:

Which Platforms Offer Multimodal Work

CrowdGen and DataAnnotation.tech both include image and audio tasks alongside text evaluation. Toloka's pivot has specifically emphasized multimodal data collection. RemoExperts includes audio and transcription projects. For specialized medical imaging work, SuperAnnotate and Labelbox (both primarily enterprise tools) have their own contractor pools, though access is more selective.

Multimodal AI training work is expected to grow faster than text-only evaluation through 2027 as AI capabilities in vision, audio, and video mature — making it a good category to develop familiarity with now rather than after the demand peak.

Which Platforms Offer Multimodal Tasks

Multimodal AI training tasks are most consistently available on Mercor (image and video evaluation tracks), DataAnnotation.tech (image description, audio transcription correction, video annotation), Alignerr (visual and audio AI evaluation), and RemoExperts (audio and speech data tasks). Outlier AI periodically opens multimodal projects. The availability is project-dependent and tends to open in waves — worth setting alerts or checking platform dashboards weekly when primary text-based tasks are slower.

Pay Premium for Multimodal Work

Multimodal evaluation tasks typically pay a premium over equivalent text tasks — image and video assessment is more cognitively demanding and requires spatial reasoning skills that text-only annotators don't have. On platforms where both text and multimodal tasks are available, experienced contractors report 15-30% higher effective hourly rates on multimodal tracks. If you have a visual, creative, or media production background, prioritise applying to multimodal tracks specifically on every platform.

Ready to Start?

Apply directly or explore our top-ranked platforms.