corpusid int64 110 292M | author int64 1.68M 2.48B | date date32 |
|---|---|---|
283,072,063 | 2,392,866,169 | 2025-11-16 |
283,073,596 | 2,202,553,944 | 2025-11-17 |
283,073,596 | 2,283,233,343 | 2025-11-17 |
283,073,596 | 2,284,687,293 | 2025-11-17 |
283,073,596 | 2,330,970,874 | 2025-11-17 |
283,073,596 | 2,355,957,162 | 2025-11-17 |
283,073,596 | 2,362,322,942 | 2025-11-17 |
283,071,945 | 31,033,341 | 2025-11-14 |
283,071,945 | 50,170,368 | 2025-11-14 |
283,071,945 | 1,404,417,826 | 2025-11-14 |
283,071,945 | 2,056,114,834 | 2025-11-14 |
283,071,945 | 2,293,723,987 | 2025-11-14 |
283,071,945 | 2,293,724,934 | 2025-11-14 |
283,071,945 | 2,294,363,357 | 2025-11-14 |
283,110,160 | 2,135,709,490 | 2025-11-20 |
283,110,160 | 2,283,189,754 | 2025-11-20 |
283,110,160 | 2,295,952,287 | 2025-11-20 |
283,110,160 | 2,302,152,165 | 2025-11-20 |
283,110,160 | 2,342,736,215 | 2025-11-20 |
283,110,160 | 2,349,949,740 | 2025-11-20 |
283,110,160 | 2,362,519,491 | 2025-11-20 |
283,110,160 | 2,367,447,542 | 2025-11-20 |
283,110,160 | 2,370,956,537 | 2025-11-20 |
283,110,160 | 2,372,143,651 | 2025-11-20 |
283,110,160 | 2,376,421,301 | 2025-11-20 |
283,110,160 | 2,393,203,260 | 2025-11-20 |
283,244,241 | 2,253,485,882 | 2025-11-23 |
283,244,241 | 2,269,463,974 | 2025-11-23 |
283,244,241 | 2,269,737,429 | 2025-11-23 |
283,244,241 | 2,325,758,086 | 2025-11-23 |
283,110,411 | 98,579,574 | 2025-11-20 |
283,110,411 | 2,208,758,300 | 2025-11-20 |
283,110,411 | 2,211,982,130 | 2025-11-20 |
283,110,411 | 2,283,250,483 | 2025-11-20 |
283,110,411 | 2,305,154,215 | 2025-11-20 |
283,449,070 | 152,806,163 | 2025-12-01 |
283,449,070 | 2,269,471,620 | 2025-12-01 |
283,449,070 | 2,353,074,569 | 2025-12-01 |
283,247,574 | 1,683,896 | 2025-12-03 |
283,247,574 | 1,700,356 | 2025-12-03 |
283,247,574 | 1,738,050 | 2025-12-03 |
283,247,574 | 1,758,085 | 2025-12-03 |
283,247,574 | 2,326,425 | 2025-12-03 |
283,247,574 | 3,106,683 | 2025-12-03 |
283,247,574 | 9,948,791 | 2025-12-03 |
283,247,574 | 40,963,426 | 2025-12-03 |
283,247,574 | 70,495,322 | 2025-12-03 |
283,247,574 | 144,038,463 | 2025-12-03 |
283,247,574 | 146,167,881 | 2025-12-03 |
283,247,574 | 2,007,538,346 | 2025-12-03 |
283,247,574 | 2,063,995,487 | 2025-12-03 |
283,247,574 | 2,100,088,196 | 2025-12-03 |
283,247,574 | 2,113,401,705 | 2025-12-03 |
283,247,574 | 2,119,244,957 | 2025-12-03 |
283,247,574 | 2,136,158,664 | 2025-12-03 |
283,247,574 | 2,138,716,089 | 2025-12-03 |
283,247,574 | 2,152,084,183 | 2025-12-03 |
283,247,574 | 2,152,802,303 | 2025-12-03 |
283,247,574 | 2,160,103,649 | 2025-12-03 |
283,247,574 | 2,187,455,323 | 2025-12-03 |
283,247,574 | 2,190,620,363 | 2025-12-03 |
283,247,574 | 2,191,406,875 | 2025-12-03 |
283,247,574 | 2,202,172,665 | 2025-12-03 |
283,247,574 | 2,211,335,875 | 2025-12-03 |
283,247,574 | 2,213,334,046 | 2025-12-03 |
283,247,574 | 2,237,634,559 | 2025-12-03 |
283,247,574 | 2,255,314,450 | 2025-12-03 |
283,247,574 | 2,258,549,455 | 2025-12-03 |
283,247,574 | 2,260,598,504 | 2025-12-03 |
283,247,574 | 2,261,672,792 | 2025-12-03 |
283,247,574 | 2,262,074,319 | 2025-12-03 |
283,247,574 | 2,267,025,893 | 2025-12-03 |
283,247,574 | 2,270,661,852 | 2025-12-03 |
283,247,574 | 2,270,664,118 | 2025-12-03 |
283,247,574 | 2,271,244,756 | 2025-12-03 |
283,247,574 | 2,283,881,714 | 2025-12-03 |
283,247,574 | 2,292,032,089 | 2025-12-03 |
283,247,574 | 2,292,034,064 | 2025-12-03 |
283,247,574 | 2,292,034,447 | 2025-12-03 |
283,247,574 | 2,309,521,685 | 2025-12-03 |
283,247,574 | 2,318,926,517 | 2025-12-03 |
283,247,574 | 2,322,096,782 | 2025-12-03 |
283,247,574 | 2,323,500,396 | 2025-12-03 |
283,247,574 | 2,324,223,442 | 2025-12-03 |
283,247,574 | 2,326,714,660 | 2025-12-03 |
283,247,574 | 2,326,842,290 | 2025-12-03 |
283,247,574 | 2,326,977,628 | 2025-12-03 |
283,247,574 | 2,330,185,774 | 2025-12-03 |
283,247,574 | 2,330,397,082 | 2025-12-03 |
283,247,574 | 2,333,408,675 | 2025-12-03 |
283,247,574 | 2,333,410,409 | 2025-12-03 |
283,247,574 | 2,333,962,294 | 2025-12-03 |
283,247,574 | 2,334,524,895 | 2025-12-03 |
283,247,574 | 2,337,789,771 | 2025-12-03 |
283,247,574 | 2,338,269,713 | 2025-12-03 |
283,247,574 | 2,338,270,113 | 2025-12-03 |
283,247,574 | 2,338,357,642 | 2025-12-03 |
283,247,574 | 2,338,444,965 | 2025-12-03 |
283,247,574 | 2,338,527,616 | 2025-12-03 |
283,247,574 | 2,343,742,870 | 2025-12-03 |
GitScholar: arXiv AI papers, their citations, and their GitHub footprint
GitScholar is a relational, fully timestamped dataset linking AI arXiv papers to their citation history on Semantic Scholar and to the GitHub repositories that reference them. It is built for studying, and predicting, how research papers gain attention over time: every row carries the date on which it became true, so the state of the whole graph can be reconstructed as of any day between 1991 and 2026-06-16.
GitScholar powers The AI Almanac, a web application that surfaces rising AI papers from their GitHub and citation activity.
Contents
The dataset is twelve Parquet files in two groups that share one key, the arXiv id.
Academic side, from arXiv metadata and the Semantic Scholar (S2) bulk dump:
| File | One row per | Rows |
|---|---|---|
features_arxiv.parquet |
arXiv paper in scope | 558,226 |
edges_author_arxiv.parquet |
(paper, author) pair for papers in scope | 2,566,255 |
edges_citation_arxiv.parquet |
(citing paper, cited paper, date) tuples, both papers in scope | 10,718,261 |
events_citation_arxiv.parquet |
(paper, date, bucket) on which a paper in scope gained citations | 16,070,452 |
edges_author_paper.parquet |
(author, paper) pair over the authors' entire publication record | 19,120,884 |
events_citation_paper.parquet |
(paper, date, bucket) for every paper in that record | 246,435,143 |
GitHub side, from a crawl of every public repository whose README mentions arXiv:
| File | One row per | Rows |
|---|---|---|
features_repo.parquet |
repository | 444,442 |
edges_repo_arxiv.parquet |
(repository, arXiv id, date) on which a README link appeared or disappeared | 3,337,278 |
events_stars.parquet |
(repository, day) on which it gained stars | 7,277,781 |
events_forks.parquet |
(repository, day) on which it gained forks | 1,814,715 |
events_issues.parquet |
(repository, day) on which issues were opened | 985,816 |
events_prs.parquet |
(repository, day) on which pull requests were opened | 900,591 |
Scope
Papers. A paper is in scope if it is on arXiv with at least one of the categories cs.AI, cs.LG, cs.CV or cs.CL, whether as primary or cross-listed category, and if Semantic Scholar holds a record for it. Papers are dated by the earlier of their first arXiv submission and their S2 publication date.
Authors. Every S2 author of a paper in scope. The author neighbourhood tables
(edges_author_paper, events_citation_paper) follow these authors across their entire
publication record, arXiv or not and in any field, so that an author's standing at a given
date can be computed. They are supersets of the two arXiv tables: semi-joining either on
features_arxiv.corpusid recovers the in-scope view.
Repositories. Every public GitHub repository whose README contained the string "arxiv" when the crawl found it, created between 2008 and the freeze. The crawl runs GitHub's search API over creation-date windows, so it is exhaustive within GitHub's own indexing of READMEs at crawl time.
Cutoff. Every table ends on 2026-06-16. Nothing dated later appears anywhere, so a point-in-time query as of any earlier day sees only what had been observed by then.
Schemas
features_arxiv
| Column | Type | Meaning |
|---|---|---|
id |
string | arXiv identifier, e.g. 2406.11190 |
corpusid |
int64 | Semantic Scholar corpus id; the key the citation tables use |
date |
date | the day the paper became public: min(first arXiv submission, S2 publication date) |
n_authors |
int32 | number of author edges the paper has in edges_author_arxiv |
Titles, abstracts and categories are not included here; they can be found in the public arXiv
metadata snapshot https://www.kaggle.com/datasets/Cornell-University/arxiv and join on id.
edges_author_arxiv, edges_author_paper
| Column | Type | Meaning |
|---|---|---|
corpusid |
int64 | the paper |
author |
int64 | Semantic Scholar author id |
date |
date | the paper's date; authorship is fixed at publication |
In edges_author_paper, papers with no S2 publication date are dated by their earliest
fully dated citation. Papers with no dated citation are omitted.
edges_citation_arxiv
| Column | Type | Meaning |
|---|---|---|
citing |
int64 | corpus id of the citing paper |
cited |
int64 | corpus id of the cited paper |
date |
date | the citing paper's date in features_arxiv |
The citation graph restricted to papers in scope: both ends are in features_arxiv, so it
can be used directly as paper-to-paper structure. It covers 507,325 citing and 379,972 cited papers.
The edge is dated by the citing paper's date as features_arxiv carries it, so that a paper
has one date throughout the dataset. The events tables below date citations by the citing
paper's S2 publication date instead; the two agree for 99.9 percent of these edges and
differ only where arXiv posted the citing paper before S2's date. Every edge has a date,
since every citing paper is in scope. Self-citations of a record by itself, an S2 merge
artifact, are removed.
events_citation_arxiv, events_citation_paper
| Column | Type | Meaning |
|---|---|---|
corpusid |
int64 | the cited paper |
date |
date, nullable | the day the citations are attributed to |
year |
int16, nullable | the year, for citations known only by year |
bucket |
enum | how the citation relates to the cited paper's date (below) |
citations |
int32 | how many citations arrived on that (date, bucket) |
A citation is dated by the publication date of the citing paper, since the citation edges themselves carry no date. Every citation in the S2 dump into a paper of the table is represented exactly once, in one of five buckets:
bucket |
date |
year |
Meaning |
|---|---|---|---|
post |
set | null | dated on or after the cited paper's date: an ordinary citation |
prepub_near |
set | null | dated up to 60 days before the cited paper's date. Usually a dating artifact: S2 defaults an unknown day to the 1st of the month, and arXiv and S2 dates differ by a median of 26 days |
prepub_far |
set | null | dated more than 60 days before the cited paper's date. The paper was public elsewhere before its arXiv posting |
year_only |
null | set | the citing paper carries a year but no day |
null |
null | null | the citing paper has neither a date nor a year |
Nothing is dropped. A consumer who wants ordinary citations
filters on bucket == "post" and can still see what that excluded. A consumer who wants
the year-only citations on a timeline chooses their own rule for placing them: a uniform
random day within the year, the year's midpoint, or the paper's own date when the years
coincide.
features_repo
| Column | Type | Meaning |
|---|---|---|
repo_id |
int32 | surrogate key used by the other GitHub tables; stable within this release only |
full_name |
string | owner/name as on GitHub; the durable identifier |
owner, name |
string | the two halves of full_name |
created |
date | GitHub's creation date for the repository |
edges_repo_arxiv
| Column | Type | Meaning |
|---|---|---|
repo_id |
int32 | the repository |
arxiv_id |
string | the paper the README links to |
date |
date | the day the link appeared (added = true) or disappeared (added = false) |
added |
bool | direction of the change |
Links are recovered from README text: arxiv.org URLs in all their forms, arXiv:NNNN.NNNNN
inline references, and eprint fields of BibTeX entries. The README is read from the
repository's commit history, so a link is dated by the commit that introduced or removed
it, collapsed to one state per day. Three consequences:
- The link set is a change log. The links a repository has on day D are the pairs whose latest event on or before D is an add. Roughly 30 percent of rows are removals; paper feed repositories in particular rotate their links daily.
arxiv_id is not restricted to papers in features_arxiv. Every arXiv link found is
kept, in any field, so the table can be joined to any arXiv-keyed dataset. Within this
release, join on features_arxiv.id to restrict to the in-scope papers.
events_stars, events_forks, events_issues, events_prs
| Column | Type | Meaning |
|---|---|---|
repo_id |
int32 | the repository |
date |
date | the day |
count |
int32 | how many stars / forks / issues / pull requests the repository gained that day |
Counts are daily gains, not running totals, and days with no gain have no row.
Stars come from GitHub's per-star starredAt stream and were never decremented for
unstars, so a repository's summed stars can slightly exceed its displayed count.
Quality filters applied to the paper set
These remove papers whose metadata is wrong in a way that would corrupt every date-based quantity derived from them.
- Author-count disagreement. Papers where arXiv and S2 disagree by more than two authors are dropped: one of the two records is typically a different paper or a corrupted merge. This is not a common case.
- Late arXiv postings. Papers with at least 25 citations, of which at least a quarter predate the paper's own date, are dropped. This catches well cited papers that were published in conferences, then uploaded to arXiv, and S2 attributes the arXiv date instead of the conference one.
- Papers with no usable author id are dropped, since they would carry an empty author
record that disagrees with
n_authors. - Papers dated before 1991-08-14, the day arXiv opened, are dropped as impossible.
Papers excluded by the first two filters are also kept out of the author neighbourhood, so that they cannot re-enter through their co-authors.
Duplicate S2 records for one arXiv id, which happen when a preprint and its published version are never merged, are resolved to the lowest corpus id: the record created first and the one citations accumulate against.
Known limitations
- Citation timing is reconstructed from the citing paper's date, so a citation from a paper S2 dates to a journal issue can appear months after the work was actually circulating, and S2's own dating errors propagate.
- The crawl finds repositories through GitHub's README search, which indexes only the default branch's README
- Link detection is textual. A repository mentioning a paper in a reading list and one implementing it produce the same edge. Repositories that link many papers can be identified from the edge table and treated separately.
- Author identity is Semantic Scholar's author disambiguation, with its known splitting and merging errors.
Sources, licensing and availability
GitScholar is available at https://huggingface.co/datasets/huawei-csl/GitScholar and is
released under the Open Data Commons Attribution License v1.0 (ODC-By), which allows reuse
and modification with appropriate attribution. The full license text is in the LICENSE
file alongside the data and at https://opendatacommons.org/licenses/by/1-0/. Attribution
should cite the paper given in the Citation section below.
It combines:
- arXiv paper metadata, obtained through the arXiv Open Archives Initiative (OAI) interface. Only identifiers and dates are redistributed here; titles, abstracts and categories remain in the public arXiv metadata and join on the arXiv id.
- Semantic Scholar paper, author and citation records from the Semantic Scholar Open Research Corpus (Kinney et al., 2023), bulk release of 2026-06-16.
- GitHub public repository metadata, collected through the official GitHub GraphQL API in accordance with GitHub's API Terms of Use. Only public metadata is included: repository names, creation dates, the arXiv identifiers found in READMEs, and daily counts of stars, forks, issues and pull requests. No personal data, no user-generated content such as README text, issue text or commit messages, and no private repository information is included.
Citation
GitScholar is described in:
GitScholar: A Dataset for Predicting AI Research Impact from GitHub Engagement. Emilien Guandalino, Lorenz K. Müller, Beatrice Alessandra Motetti, Konstantin Berestizshevsky and Lukas Cavigelli. arXiv:2609.26361, 2026. https://arxiv.org/abs/2609.26361
@article{guandalino2026gitscholar,
title = {GitScholar: A Dataset for Predicting AI Research Impact from GitHub Engagement},
author = {Guandalino, Emilien and Müller, Lorenz K. and Motetti, Beatrice Alessandra
and Berestizshevsky, Konstantin and Cavigelli, Lukas},
journal = {arXiv preprint arXiv:2609.26361},
year = {2026},
url = {https://arxiv.org/abs/2609.26361}
}
GitScholar is a derived database of the Semantic Scholar Academic Graph, whose ODC-By attribution requirement carries over. Please also cite:
The Semantic Scholar Open Data Platform. Rodney Kinney, Chloe Anastasiades, Russell Authur, Iz Beltagy, Jonathan Bragg, Alexandra Buraczynski, Isabel Cachola, Stefan Candra, Yoganand Chandrasekhar, Arman Cohan, et al. arXiv:2301.10140, 2023. https://arxiv.org/abs/2301.10140
@article{kinney2023semantic,
title = {The Semantic Scholar Open Data Platform},
author = {Kinney, Rodney and Anastasiades, Chloe and Authur, Russell and Beltagy, Iz and
Bragg, Jonathan and Buraczynski, Alexandra and Cachola, Isabel and Candra, Stefan
and Chandrasekhar, Yoganand and Cohan, Arman and others},
journal = {arXiv preprint arXiv:2301.10140},
year = {2023},
url = {https://arxiv.org/abs/2301.10140}
}
Reading the data
All files are standard Parquet and open with any Parquet reader. In Python:
import polars as pl
papers = pl.read_parquet("features_arxiv.parquet")
events = pl.read_parquet("events_citation_arxiv.parquet")
# Citations at one year after publication, ordinary citations only
cit_1y = (
events.filter(pl.col("bucket") == "post")
.join(papers.select("corpusid", pl.col("date").alias("pub_date")), on="corpusid")
.filter(pl.col("date") <= pl.col("pub_date").dt.offset_by("1y"))
.group_by("corpusid").agg(pl.col("citations").sum())
)
# Stars of every repository linking a paper, as of a cutoff date
cutoff = pl.date(2025, 6, 17)
links = (
pl.read_parquet("edges_repo_arxiv.parquet")
.filter(pl.col("date") <= cutoff).sort("date")
.group_by("repo_id", "arxiv_id").agg(pl.col("added").last())
.filter("added")
)
stars = (
pl.read_parquet("events_stars.parquet")
.filter(pl.col("date") <= cutoff)
.group_by("repo_id").agg(pl.col("count").sum().alias("stars"))
)
stars_per_paper = links.join(stars, on="repo_id", how="left").group_by("arxiv_id").agg(pl.col("stars").sum())
- Downloads last month
- 411