‘Captchas’: Has Google got us training its AI without knowing it?
The service was created 15 years ago with a dual purpose: as a security filter against bots and as a generator of datasets to train artificial intelligence
Anyone who browses the internet is familiar with captchas. They are small tests to prove we are human and not a bot. They can be a set of images where you have to click, for example, on those that show a bus or a tractor; sometimes it’s a puzzle piece you must place in its slot; other times you are simply asked to check a box that says “I’m not a robot.”
Several companies provide captcha services that prevent a website from being overwhelmed by bot traffic. But the main supplier of these puzzles is Google. In recent months the company has come under scrutiny for this reason. The triggers were a viral tweet and a change in Google’s terms for captchas that shifted its role from data controller to data processor.
The tweet (now “from a suspended account”) set off a flood of posts on social media and blogs describing how Google used data generated by users solving captchas to train its AI. For example, selecting buses in multiple images helps label them accurately so the software can learn what a bus looks like. That would be useful for Waymo, Alphabet’s self-driving car unit, which is also Google’s parent company. But to what extent is that actually happening?
The claim has circulated online for years. Added to that are privacy concerns from bodies such as France’s data protection regulator (CNIL) and the Office for Data Protection Supervision of the State of Bavaria in Germany. The first argued that the purpose of captchas is not solely security, and the second raised doubts about transparency in processing this information.
Given the value of data generated by captchas, would it make sense to use it to train AI models? “Most [captchas] are image datasets related to traffic,” says Karina Gibert, professor at the Universitat Politècnica de Catalunya (UPC), and former director and founder of the Intelligent Data Science and Artificial Intelligence Research Center. “They are closely linked to recognizing obstacles, vehicles, and traffic signs in motion. That connects a lot with what an autonomous car needs to see, understand, and interpret: it has to operate in an environment where it will encounter other vehicles or traffic signs it has to interpret.”
Object recognition is one of the essential factors in developing autonomous cars. “When you deploy a self-driving system in a real environment, where the camera can simply be dirty or it’s raining, having images where a human has labeled a blurry tractor helps the system improve its ability to detect tractors,” explains Jordi Nin, a researcher at the Institute for Artificial Intelligence Research. “What the system does is try to recognize the pattern the human used to label ‘tractor.’” People can be wrong sometimes, but when doing this kind of labeling an image is typically shown to several users. If many indicate a tractor is present, the label is accepted.
Labeling blurry images correctly or those where the object is poorly perceived helps handle exceptions more precisely. “You need a very large database with many images that include traffic lights and that establish where there is a traffic light and where there isn’t,” Gibert says. “With this immense database, you train the algorithm and it becomes capable of recognizing a traffic light in any position, from any angle, with clouds, strong sun, rain, or even in the jungle.”
The fate of captcha data
One of the key figures in the birth of captchas was Guatemalan Luis von Ahn, who later co-founded Duolingo. In the 2000s he created the reCAPTCHA service, whose original purpose was to protect web pages using tests that then-circa programs could not solve. That filtered out bots, which were already roaming the internet. The captcha system also had another goal: digitizing old books. The puzzles consisted of written words that, once correctly labeled by users, helped an algorithm decipher texts better. Those were still days of digital naivety, with the glow of pure experimentation and no profit calculations.
Google acquired reCAPTCHA in 2009, expanded the service and used it to digitize millions of books and newspapers. That task involved training algorithms to read all kinds of texts with increasing accuracy. For years the company clearly stated it used captchas to strengthen its artificial intelligence systems. In a previous version of the service’s official page the slogan read “Stop a bot, build a bot.” It continued: “High-quality, human-labeled images are gathered into datasets that can be used to train machine learning systems.”
According to Google Spain, that is no longer the case: “In 2018, the rapid advance of AI and the changing cyberthreat landscape marked a turning point. With the launch of reCAPTCHA v3 we stopped using visual challenges and we no longer use reCAPTCHA challenge responses to train models.” Today, any reference to model training no longer appears on the service’s official page and the company says data collection is carried out only to improve the captchas.
However, some doubts persist. For starters, many websites still use an older version of the service, which did serve to train AI. Also, reCAPTCHA is divided into two products: one for businesses, whose terms clearly specify that collected data is used only to bolster security, and a free version. The free version was governed until last April by Google’s general privacy policy and terms of service. Some sites still display that reference. And the privacy terms stipulated that user data could be used to improve Google services or develop new ones. The door was left open.
In 2023, a paper from the University of California, Irvine suggested that Google was using its captchas (the version prior to v3) to obtain labeled images and training data for its AI. Some experts side with the researchers who authored that paper. “We assume Google is using them to train algorithms that recognize elements in the traffic domain,” Gibert says.
“To train a computer vision system you have to use supervised learning algorithms that work on labeled datasets,” Gibert explains. “To train them you need lots of labeled data. And with captchas you generate a human-labeled dataset of extremely high quality that keeps growing.”
Specialists agree that data generated by captchas would be highly valuable. Nin emphasizes that quality training data has become scarce and highly sought-after: “Google gives you its model training libraries, even provides small instances semi-free so you can do training, tutorials and training courses. But the asset needed to build a model of the quality Google requires is the data. As long as they control the data, you don’t compete with them.”
Google’s change of role
On April 2 Google officially changed its role in the reCAPTCHA service. It moved from data controller to data processor. Acting as a processor, Google can no longer unilaterally decide how to use the information collected. It would need to rely on the websites that use the service, which become the data controllers. In addition, the provision of the service will be governed by Google Cloud Platform terms, which are more restrictive than the general privacy policy and terms of use.
Gibert highlights another aspect: “The most important implication is that the user distributes the role of data controller,” she says. “The website that integrates a captcha is the one that now has controller responsibilities, while Google can collect information and analyze it. They have only shifted liability for misuse of the data to the person who places the captcha on their website.” It seems uncertainties are not over yet.
Sign up for our weekly newsletter to get more English-language news coverage from EL PAÍS USA Edition