Curated datasets designed for training artificial intelligence and large language models.
~$25,000/yr est.
On-demand
1
1
Freshness
Recently enriched
Complete
75%
API
API Available
Get an instant price estimate based on your organization profile — seats, usage rights, and contract term.Estimated ~$25,000/yr est.
Request a sample directly from Euromonitor International.
[2402.18041] Datasets for Large Language Models
by Y Liu · 2024 · Cited by 234 — This paper embarks on an exploration into the Large Language Model (LLM) datasets, which play a crucial role in the remarkable advancements of LLMs.
RedPajama: an Open Dataset for Training Large ...
Together, the RedPajama datasets comprise over 100 trillion tokens spanning multiple domains and with their quality signals facilitate the filtering of data, ...
The Largest Collection of Ethical Data for LLM Pre-Training
by PC Langlais · 2025 · Cited by 12 — In this paper, we introduce Common Corpus, the largest open dataset for language model pre-training. The data assembled in Common Corpus are ...
Expected fields and columns in this data product
AI/Large Language Model (LLM) Training Data is an alternative data product offered by Euromonitor International, available on discovery. Data is updated on-demand. API access is available for programmatic integration. Vedex estimates pricing at roughly $25,000/yr (an estimate, not a vendor-published price).
Curated datasets designed for training artificial intelligence and large language models.
Euromonitor International is a data provider vendor based in London, England. Euromonitor International is a leading global provider of strategic market research, offering comprehensive data, analysis, and insights on industries, economies, and consumers worldwide to help busin...