March 2010 April 2010 May 2010 June 2010 July 2010
August 2010
September 2010 October 2010 November 2010 December 2010 January 2011 February 2011 March 2011 April 2011 May 2011 June 2011 July 2011 August 2011 September 2011 October 2011 November 2011 December 2011 January 2012 February 2012 March 2012 April 2012 May 2012 June 2012 July 2012 August 2012 September 2012 October 2012 November 2012 December 2012 January 2013 February 2013 March 2013 April 2013 May 2013 June 2013 July 2013 August 2013 September 2013 October 2013 November 2013 December 2013 January 2014 February 2014 March 2014 April 2014 May 2014 June 2014 July 2014 August 2014 September 2014 October 2014 November 2014 December 2014 January 2015 February 2015 March 2015 April 2015 May 2015 June 2015 July 2015 August 2015 September 2015 October 2015 November 2015 December 2015 January 2016 February 2016 March 2016 April 2016 May 2016 June 2016 July 2016 August 2016 September 2016 October 2016 November 2016 December 2016 January 2017 February 2017 March 2017 April 2017 May 2017 June 2017 July 2017 August 2017 September 2017 October 2017 November 2017 December 2017 January 2018 February 2018 March 2018 April 2018 May 2018 June 2018 July 2018 August 2018 September 2018 October 2018 November 2018 December 2018 January 2019 February 2019 March 2019 April 2019 May 2019 June 2019 July 2019 August 2019 September 2019 October 2019 November 2019 December 2019 January 2020 February 2020 March 2020 April 2020 May 2020 June 2020 July 2020 August 2020 September 2020 October 2020 November 2020 December 2020 January 2021 February 2021 March 2021 April 2021 May 2021 June 2021 July 2021 August 2021 September 2021 October 2021 November 2021 December 2021 January 2022 February 2022 March 2022 April 2022 May 2022 June 2022 July 2022 August 2022 September 2022 October 2022 November 2022 December 2022 January 2023 February 2023 March 2023 April 2023 May 2023 June 2023 July 2023 August 2023 September 2023 October 2023 November 2023 December 2023 January 2024 February 2024 March 2024 April 2024 May 2024 June 2024 July 2024 August 2024 September 2024 October 2024 November 2024 December 2024
1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18
19
20
21
22
23
24
25
26
27
28
29
30
31
News Every Day |

This is where the data to build AI comes from

AI is all about data. Reams and reams of data are needed to train algorithms to do what we want, and what goes into the AI models determines what comes out. But here’s the problem: AI developers and researchers don’t really know much about the sources of the data they are using. AI’s data collection practices are immature compared with the sophistication of AI model development. Massive data sets often lack clear information about what is in them and where it came from. 

The Data Provenance Initiative, a group of over 50 researchers from both academia and industry, wanted to fix that. They wanted to know, very simply: Where does the data to build AI come from? They audited nearly 4,000 public data sets spanning over 600 languages, 67 countries, and three decades. The data came from 800 unique sources and nearly 700 organizations. 

Their findings, shared exclusively with MIT Technology Review, show a worrying trend: AI’s data practices risk concentrating power overwhelmingly in the hands of a few dominant technology companies. 

In the early 2010s, data sets came from a variety of sources, says Shayne Longpre, a researcher at MIT who is part of the project. 

It came not just from encyclopedias and the web, but also from sources such as parliamentary transcripts, earning calls, and weather reports. Back then, AI data sets were specifically curated and collected from different sources to suit individual tasks, Longpre says.

Then transformers, the architecture underpinning language models, were invented in 2017, and the AI sector started seeing performance get better the bigger the models and data sets were. Today, most AI data sets are built by indiscriminately hoovering material from the internet. Since 2018, the web has been the dominant source for data sets used in all media, such as audio, images, and video, and a gap between scraped data and more curated data sets has emerged and widened.

“In foundation model development, nothing seems to matter more for the capabilities than the scale and heterogeneity of the data and the web,” says Longpre. The need for scale has also boosted the use of synthetic data massively.

The past few years have also seen the rise of multimodal generative AI models, which can generate videos and images. Like large language models, they need as much data as possible, and the best source for that has become YouTube. 

For video models, as you can see in this chart, over 70% of data for both speech and image data sets comes from one source.

This could be a boon for Alphabet, Google’s parent company, which owns YouTube. Whereas text is distributed across the web and controlled by many different websites and platforms, video data is extremely concentrated in one platform.

“It gives a huge concentration of power over a lot of the most important data on the web to one company,” says Longpre. 

And because Google is also developing its own AI models, its massive advantage also raises questions about how the company will make this data available for competitors, says Sarah Myers West, the co–executive director at the AI Now Institute.

“It’s important to think about data not as though it’s sort of this naturally occurring resource, but it’s something that is created through particular processes,” says Myers West.

“If the data sets on which most of the AI that we’re interacting with reflect the intentions and the design of big, profit-motivated corporations—that’s reshaping the infrastructures of our world in ways that reflect the interests of those big corporations,” she says.

This monoculture also raises questions about how accurately the human experience is portrayed in the data set and what kinds of models we are building, says Sara Hooker, the vice president of research at the technology company Cohere, who is also part of the Data Provenance Initiative.

People upload videos to YouTube with a particular audience in mind, and the way people act in those videos is often intended for very specific effect. “Does [the data] capture all the nuances of humanity and all the ways that we exist?” says Hooker. 

Hidden restrictions

AI companies don’t usually share what data they used to train their models. One reason is that they want to protect their competitive edge. The other is that because of the complicated and opaque way data sets are bundled, packaged, and distributed, they likely don’t even know where all the data came from.

They also probably don’t have complete information about any constraints on how that data is supposed to be used or shared. The researchers at the Data Provenance Initiative found that data sets often have restrictive licenses or terms attached to them, which should limit their use for commercial purposes, for example.

“This lack of consistency across the data lineage makes it very hard for developers to make the right choice about what data to use,” says Hooker.

It also makes it almost impossible to be completely certain you haven’t trained your model on copyrighted data, adds Longpre.

More recently, companies such as OpenAI and Google have struck exclusive data-sharing deals with publishers, major forums such as Reddit, and social media platforms on the web. But this becomes another way for them to concentrate their power.

“These exclusive contracts can partition the internet into various zones of who can get access to it and who can’t,” says Longpre.

The trend benefits the biggest AI players, who can afford such deals, at the expense of researchers, nonprofits, and smaller companies, who will struggle to get access. The largest companies also have the best resources for crawling data sets.

“This is a new wave of asymmetric access that we haven’t seen to this extent on the open web,” Longpre says.

The West vs. the rest

The data that is used to train AI models is also heavily skewed to the Western world. Over 90% of the data sets that the researchers analyzed came from Europe and North America, and fewer than 4% came from Africa. 

“These data sets are reflecting one part of our world and our culture, but completely omitting others,” says Hooker.

The dominance of the English language in training data is partly explained by the fact that the internet is still over 90% in English, and there are still a lot of places on Earth where there’s really poor internet connection or none at all, says Giada Pistilli, principal ethicist at Hugging Face, who was not part of the research team. But another reason is convenience, she adds: Putting together data sets in other languages and taking other cultures into account requires conscious intention and a lot of work. 

The Western focus of these data sets becomes particularly clear with multimodal models. When an AI model is prompted for the sights and sounds of a wedding, for example, it might only be able to represent Western weddings, because that’s all that it has been trained on, Hooker says. 

This reinforces biases and could lead to AI models that push a certain US-centric worldview, erasing other languages and cultures.

“We are using these models all over the world, and there’s a massive discrepancy between the world we’re seeing and what’s invisible to these models,” Hooker says. 

Москва

Экс-главу депкультуры Москвы доставили в суд в футболке с надписью об СВО

Bumrah bowls a double wicket over for 12th five-wicket haul

Steve Smith breaks Steve Waugh's record

French mass rape trial adjourns ahead of verdict expected on December 19

D Gukesh: World champion shares life lessons by his mother

Ria.city






Read also

India’s economic growth slowdown is ‘temporary blip’ – New Delhi

Father of teenager who died from poisoned alcohol speaks out for the first time

‘I’m not paying that amount to have a jobby’ joke Scottish shoppers after spotting item’s hilarious name in John Lewis

News, articles, comments, with a minute-by-minute update, now on Today24.pro

News Every Day

Steve Smith breaks Steve Waugh's record

Today24.pro — latest news 24/7. You can add your news instantly now — here


News Every Day

School and road closures in Manitoba on Monday



Sports today


Новости тенниса
Янник Синнер

В. Березуцкий вспомнил, как в Китае его перепутали с первой ракеткой мира: «Орут «Синнер», я машу. Ревели, кричали»



Спорт в России и мире
Москва

Спортсмены из Росгвардии заняли призовые места на этапе Кубка России по лыжным гонкам в Пермском крае



All sports news today





Sports in Russia today

Москва

«Русская классика» в Туле: АКМ сыграет с «Рубином» под открытым небом


Новости России

Game News

Critical metals for electronic components and gadgets jump in price, as China's trade restrictions with the US begin to bite


Russian.city


Москва

«То пандемия, то война»: в паломнических службах рассказали, почему так мало россиян теперь ездят к Николаю Чудотворцу


Губернаторы России
Александр Большунов

Серебро Большунова, доминирование Сливко и провал Бажина: чем завершились третьи этапы КР по лыжным гонкам и биатлону


"Поедет домой": На похороны Яниса Тиммы собирают деньги – Седокова не участвует

Щелкунчик с Фарухом Рузиматовым на сцене Александринского театра Санкт- Петербурга  

Рилсмейкер. Услуги Рилсмейкера.

Гидрометцентр: В Москве в пятницу начнется оттепель до плюс трех


Завершение строительства трамвайной линии «Славянка» сдвинули на 2026 г.

Алсу отменила концерт в Москве из-за проблем со здоровьем

Девушка рэпера Тимати Валентина выложила фото в полотенце

Масспостинг вертикальных видео в TikTok, Youtube-shorts, ВК-клипы, Reels.


Эрика Андреева проиграла в финале турнира WTA 125 в Лиможе в парном разряде

Арина Соболенко выложила эффектные фото в коротком платье

Тренер второй ракетки мира сделал признание о Елене Рыбакиной

Маннарино: руководство хорошо постаралось, чтобы выставить Синнера и Свёнтек жертвами в допинг-делах



Дмитрий Миляев провел заседание рабочей группы по подготовке Совета при президенте РФ

Как проходит процедура интимного омоложения

Рилсмейкер. Услуги Рилсмейкера.

Видеопоздравление от Деда Мороза и детективов агентства «ХРУМ»


Как проходит процедура интимного омоложения

Мурманчане представят заполярную столицу на Чемпионате России по самбо

Рилсмейкер. Услуги Рилсмейкера.

Бойцы из Владимирской области одержали победы на турнире ММА в Лужниках


Пензенские журналисты прибыли на пресс-конференцию Владимира Путина

Русские учёные нашли лекарство, которое продлит питомцам жизнь

Московский НПЗ получил престижную награду Лидер качества

Shot: Неизвестный кинул бутылку с горючим в полицейскую машину в Москве



Путин в России и мире






Персональные новости Russian.city
Земфира

Танылган блогер Земфира Солтанова үлгән



News Every Day

D Gukesh: World champion shares life lessons by his mother




Friends of Today24

Музыкальные новости

Персональные новости