

If you run Nemotron, from open datasets.
See the dataset table here. There are sections for pretraining and for specific tasks/topics, like STEM, code and such: https://huggingface.co/nvidia
Chinese models are starting to open their pretraining datasets too, albeit slowly.


Admittedly, no, but this is slowly improving. Stepfun published their SFT dataset, smaller labs are publishing their datasets for task specific tunes. I believe there was another Chinese lab that published bulk pretraining data, but I can’t find it in my history at the moment.
And, notably, these comparatively tiny labs generally aren’t scraping the internet so abusively like OpenAI/Meta. They don’t need as much. Going by statements in their papers, they tend to use existing archives of web data, buy commercial data, or (more recently) generate a lot synthetic data.