AI and Society

Who Trains the Trainers? Inside the AI Shadow Workforce

AI shadow workforce illustration of hidden workers behind glowing AI interface
Who Trains the Trainers

By Stuart Kerr, Technology Correspondent, LiveAIWire

The AI shadow workforce is what makes every polished AI product possible, and it remains almost entirely invisible in how the technology is discussed. Every time an AI system correctly identifies a pedestrian in a camera feed, accurately transcribes a difficult accent, or appropriately refuses a harmful request, a human being made that possible. This work is foundational to almost everything the technology can do, and it is scattered across dozens of countries, working through platforms that connect task supply to task demand at industrial scale.

LiveAIWire has already documented the wage conditions this workforce faces in our reporting on low-wage clickworkers training AI systems. What deserves closer attention is a specific and consequential slice of that work: the human feedback that shapes how AI models actually behave, and the early efforts by these workers to organise for better conditions.

The RLHF Problem: Who Are the Humans in the Loop?

One of the most consequential forms of work within the AI shadow workforce is the human feedback that trains large language models through reinforcement learning from human feedback. Human raters presented with pairs of AI outputs judge which is better across dimensions including accuracy, helpfulness, and safety. Their judgments reflect their values, their cultural context, and the economic pressures they operate under.

The Algorithmic Justice League and others have documented how cultural and linguistic assumptions embedded in training data and human feedback shape AI system outputs in ways that can disadvantage users from different cultural contexts. These are not random errors, they are systematic patterns reflecting whose judgments were used to train the system and whose were not.

The Global Geography of the AI Shadow Workforce

AI data work is distributed across the globe in a pattern reflecting labour cost differentials and the availability of specific language and cultural competencies. Countries including Kenya, the Philippines, India, Venezuela, and Pakistan are significant hubs. Platforms including Scale AI, Appen, and Sama connect this workforce to technology company clients. The International Labour Organisation’s research on digital labour platforms documents the precarious nature of employment: tasks allocated algorithmically, earnings fluctuating unpredictably, access to work suspendable without appeal, and most workers having no access to social protection including sick pay, maternity leave, or pensions.

Organising the AI Shadow Workforce

The AI shadow workforce is not passive. In Kenya, Uganda, and the Philippines, workers at data annotation facilities have formed associations, staged work stoppages, and engaged in legal proceedings to challenge their employment terms. In the United States, content moderators at companies including Cognizant and Accenture have organised through unions and filed lawsuits alleging inadequate mental health support.

Technology companies have responded with improved support services and increased pay at facilities under public scrutiny, but critics argue these measures are insufficient relative to the scale of the workforce and the revenues generated by the systems their labour makes possible. This dynamic echoes what LiveAIWire has found in our coverage of who really controls the gig economy, where algorithmic management concentrates decision-making power with platforms while distributing risk to the workers subject to them.

Toward Visibility and Accountability

The EU AI Act includes requirements for transparency about training data, and several European countries are developing platform work directives that would extend employee protections to gig workers. The question of how AI development is paid for, in labour as well as capital, is one of the defining ethical questions of the technology’s current phase. As LiveAIWire has examined in our coverage of AI and data equity in agriculture, communities that generate the foundational data on which commercial AI depends deserve governance frameworks that recognise their contribution, whether that data comes from a smallholder farm or a content moderation queue.

What Would Fair Treatment of the AI Shadow Workforce Look Like?

The contrast between the economic value generated by AI systems and the labour conditions of the AI shadow workforce that enables them has prompted debate about what fair AI labour practices would actually require. Researchers, labour organisations, and some technology companies have proposed frameworks that include minimum pay floors above those dictated by local market conditions, portable benefits for workers whose employment is contingent and variable, limits on the psychological harm associated with content moderation work, and transparency about how worker data is used by the platforms that employ them.

Some technology companies have made voluntary commitments in these areas. Google, Microsoft, and others have published supplier codes of conduct that specify requirements for data annotation vendors, including provisions on pay, working hours, and mental health support. The enforceability of these commitments across complex global supply chains is limited, and the gap between published standards and working conditions on the ground, documented repeatedly by investigative journalism, suggests that voluntary commitments are not sufficient without external accountability.

About the Author

Stuart Kerr is Technology Correspondent at LiveAIWire, covering artificial intelligence, cybersecurity, and the social impact of emerging technology. He publishes daily at LiveAIWire.com.