---
format: "aidr-story-markdown/v1"
id: "8dcd44f75536cbb4f4c72c55919bab4c7d77ccf1c32056118a7559ee3eb36234"
canonical_url: "https://aidr.today/8dcd44f7?lang=en"
title: "User Releases First Complete 16 TB arXiv Corpus of 3.1 Million Papers on Hugging Face"
lang: "en"
requested_lang: "en"
available_langs: ["en","vi"]
translation_fallback: null
fallback_fields: []
published_at: "2026-09-20T05:36:58.000Z"
category: "Infra"
topics: ["arxiv","huggingface","dataset","open-source","academic","corpus","latex","pdf"]
source_urls: ["https://marketbrief.now/ai/user-releases-first-complete-16-tb-arxiv-corpus-of-31-million-papers-on-1cb9eb7b","https://x.com/lhoestq/status/2101432725085241801","https://x.com/eliebakouch/status/2101426116892266820","https://x.com/TheZachMueller/status/2101421105198047301","https://x.com/secemp9/status/2101416879772340411","https://huggingnews.com/ai/user-releases-first-complete-16-tb-arxiv-corpus-of-31-million-papers-on-1cb9eb7b"]
summary: "An extensive new archive of academic literature is now available as an open dataset on the Hugging Face platform. The 16 TB collection, shared by user @secemp9…"
---

# User Releases First Complete 16 TB arXiv Corpus of 3\.1 Million Papers on Hugging Face

> [Open the canonical story](<https://aidr.today/8dcd44f7?lang=en>)

**Published:** 2026-09-20T05:36:58.000Z
**Category:** Infra
**Topics:** arxiv, huggingface, dataset, open\-source, academic, corpus, latex, pdf

## Summary

An extensive new archive of academic literature is now available as an open dataset on the Hugging Face platform\. The 16 TB collection, shared by user @secemp9…

## Sources

- [Story source](<https://marketbrief.now/ai/user-releases-first-complete-16-tb-arxiv-corpus-of-31-million-papers-on-1cb9eb7b>)
- [Story source](<https://x.com/lhoestq/status/2101432725085241801>)
- [Supporting source](<https://x.com/eliebakouch/status/2101426116892266820>)
- [Supporting source](<https://x.com/TheZachMueller/status/2101421105198047301>)
- [Story source](<https://x.com/secemp9/status/2101416879772340411>)
- [Story source](<https://huggingnews.com/ai/user-releases-first-complete-16-tb-arxiv-corpus-of-31-million-papers-on-1cb9eb7b>)

