Datasets:

Tasks:
Other
Modalities:
Tabular
Text
Formats:
parquet
Languages:
English
ArXiv:
License:
Dataset Viewer
Auto-converted to Parquet Duplicate
Search is not available for this dataset
corpusid
int64
110
292M
author
int64
1.68M
2.48B
date
date32
283,072,063
2,392,866,169
2025-11-16
283,073,596
2,202,553,944
2025-11-17
283,073,596
2,283,233,343
2025-11-17
283,073,596
2,284,687,293
2025-11-17
283,073,596
2,330,970,874
2025-11-17
283,073,596
2,355,957,162
2025-11-17
283,073,596
2,362,322,942
2025-11-17
283,071,945
31,033,341
2025-11-14
283,071,945
50,170,368
2025-11-14
283,071,945
1,404,417,826
2025-11-14
283,071,945
2,056,114,834
2025-11-14
283,071,945
2,293,723,987
2025-11-14
283,071,945
2,293,724,934
2025-11-14
283,071,945
2,294,363,357
2025-11-14
283,110,160
2,135,709,490
2025-11-20
283,110,160
2,283,189,754
2025-11-20
283,110,160
2,295,952,287
2025-11-20
283,110,160
2,302,152,165
2025-11-20
283,110,160
2,342,736,215
2025-11-20
283,110,160
2,349,949,740
2025-11-20
283,110,160
2,362,519,491
2025-11-20
283,110,160
2,367,447,542
2025-11-20
283,110,160
2,370,956,537
2025-11-20
283,110,160
2,372,143,651
2025-11-20
283,110,160
2,376,421,301
2025-11-20
283,110,160
2,393,203,260
2025-11-20
283,244,241
2,253,485,882
2025-11-23
283,244,241
2,269,463,974
2025-11-23
283,244,241
2,269,737,429
2025-11-23
283,244,241
2,325,758,086
2025-11-23
283,110,411
98,579,574
2025-11-20
283,110,411
2,208,758,300
2025-11-20
283,110,411
2,211,982,130
2025-11-20
283,110,411
2,283,250,483
2025-11-20
283,110,411
2,305,154,215
2025-11-20
283,449,070
152,806,163
2025-12-01
283,449,070
2,269,471,620
2025-12-01
283,449,070
2,353,074,569
2025-12-01
283,247,574
1,683,896
2025-12-03
283,247,574
1,700,356
2025-12-03
283,247,574
1,738,050
2025-12-03
283,247,574
1,758,085
2025-12-03
283,247,574
2,326,425
2025-12-03
283,247,574
3,106,683
2025-12-03
283,247,574
9,948,791
2025-12-03
283,247,574
40,963,426
2025-12-03
283,247,574
70,495,322
2025-12-03
283,247,574
144,038,463
2025-12-03
283,247,574
146,167,881
2025-12-03
283,247,574
2,007,538,346
2025-12-03
283,247,574
2,063,995,487
2025-12-03
283,247,574
2,100,088,196
2025-12-03
283,247,574
2,113,401,705
2025-12-03
283,247,574
2,119,244,957
2025-12-03
283,247,574
2,136,158,664
2025-12-03
283,247,574
2,138,716,089
2025-12-03
283,247,574
2,152,084,183
2025-12-03
283,247,574
2,152,802,303
2025-12-03
283,247,574
2,160,103,649
2025-12-03
283,247,574
2,187,455,323
2025-12-03
283,247,574
2,190,620,363
2025-12-03
283,247,574
2,191,406,875
2025-12-03
283,247,574
2,202,172,665
2025-12-03
283,247,574
2,211,335,875
2025-12-03
283,247,574
2,213,334,046
2025-12-03
283,247,574
2,237,634,559
2025-12-03
283,247,574
2,255,314,450
2025-12-03
283,247,574
2,258,549,455
2025-12-03
283,247,574
2,260,598,504
2025-12-03
283,247,574
2,261,672,792
2025-12-03
283,247,574
2,262,074,319
2025-12-03
283,247,574
2,267,025,893
2025-12-03
283,247,574
2,270,661,852
2025-12-03
283,247,574
2,270,664,118
2025-12-03
283,247,574
2,271,244,756
2025-12-03
283,247,574
2,283,881,714
2025-12-03
283,247,574
2,292,032,089
2025-12-03
283,247,574
2,292,034,064
2025-12-03
283,247,574
2,292,034,447
2025-12-03
283,247,574
2,309,521,685
2025-12-03
283,247,574
2,318,926,517
2025-12-03
283,247,574
2,322,096,782
2025-12-03
283,247,574
2,323,500,396
2025-12-03
283,247,574
2,324,223,442
2025-12-03
283,247,574
2,326,714,660
2025-12-03
283,247,574
2,326,842,290
2025-12-03
283,247,574
2,326,977,628
2025-12-03
283,247,574
2,330,185,774
2025-12-03
283,247,574
2,330,397,082
2025-12-03
283,247,574
2,333,408,675
2025-12-03
283,247,574
2,333,410,409
2025-12-03
283,247,574
2,333,962,294
2025-12-03
283,247,574
2,334,524,895
2025-12-03
283,247,574
2,337,789,771
2025-12-03
283,247,574
2,338,269,713
2025-12-03
283,247,574
2,338,270,113
2025-12-03
283,247,574
2,338,357,642
2025-12-03
283,247,574
2,338,444,965
2025-12-03
283,247,574
2,338,527,616
2025-12-03
283,247,574
2,343,742,870
2025-12-03
End of preview. Expand in Data Studio

GitScholar: arXiv AI papers, their citations, and their GitHub footprint

GitScholar is a relational, fully timestamped dataset linking AI arXiv papers to their citation history on Semantic Scholar and to the GitHub repositories that reference them. It is built for studying, and predicting, how research papers gain attention over time: every row carries the date on which it became true, so the state of the whole graph can be reconstructed as of any day between 1991 and 2026-06-16.

GitScholar powers The AI Almanac, a web application that surfaces rising AI papers from their GitHub and citation activity.

The AI Almanac

Contents

The dataset is twelve Parquet files in two groups that share one key, the arXiv id.

Academic side, from arXiv metadata and the Semantic Scholar (S2) bulk dump:

File One row per Rows
features_arxiv.parquet arXiv paper in scope 558,226
edges_author_arxiv.parquet (paper, author) pair for papers in scope 2,566,255
edges_citation_arxiv.parquet (citing paper, cited paper, date) tuples, both papers in scope 10,718,261
events_citation_arxiv.parquet (paper, date, bucket) on which a paper in scope gained citations 16,070,452
edges_author_paper.parquet (author, paper) pair over the authors' entire publication record 19,120,884
events_citation_paper.parquet (paper, date, bucket) for every paper in that record 246,435,143

GitHub side, from a crawl of every public repository whose README mentions arXiv:

File One row per Rows
features_repo.parquet repository 444,442
edges_repo_arxiv.parquet (repository, arXiv id, date) on which a README link appeared or disappeared 3,337,278
events_stars.parquet (repository, day) on which it gained stars 7,277,781
events_forks.parquet (repository, day) on which it gained forks 1,814,715
events_issues.parquet (repository, day) on which issues were opened 985,816
events_prs.parquet (repository, day) on which pull requests were opened 900,591

Scope

Papers. A paper is in scope if it is on arXiv with at least one of the categories cs.AI, cs.LG, cs.CV or cs.CL, whether as primary or cross-listed category, and if Semantic Scholar holds a record for it. Papers are dated by the earlier of their first arXiv submission and their S2 publication date.

Authors. Every S2 author of a paper in scope. The author neighbourhood tables (edges_author_paper, events_citation_paper) follow these authors across their entire publication record, arXiv or not and in any field, so that an author's standing at a given date can be computed. They are supersets of the two arXiv tables: semi-joining either on features_arxiv.corpusid recovers the in-scope view.

Repositories. Every public GitHub repository whose README contained the string "arxiv" when the crawl found it, created between 2008 and the freeze. The crawl runs GitHub's search API over creation-date windows, so it is exhaustive within GitHub's own indexing of READMEs at crawl time.

Cutoff. Every table ends on 2026-06-16. Nothing dated later appears anywhere, so a point-in-time query as of any earlier day sees only what had been observed by then.

Schemas

features_arxiv

Column Type Meaning
id string arXiv identifier, e.g. 2406.11190
corpusid int64 Semantic Scholar corpus id; the key the citation tables use
date date the day the paper became public: min(first arXiv submission, S2 publication date)
n_authors int32 number of author edges the paper has in edges_author_arxiv

Titles, abstracts and categories are not included here; they can be found in the public arXiv metadata snapshot https://www.kaggle.com/datasets/Cornell-University/arxiv and join on id.

edges_author_arxiv, edges_author_paper

Column Type Meaning
corpusid int64 the paper
author int64 Semantic Scholar author id
date date the paper's date; authorship is fixed at publication

In edges_author_paper, papers with no S2 publication date are dated by their earliest fully dated citation. Papers with no dated citation are omitted.

edges_citation_arxiv

Column Type Meaning
citing int64 corpus id of the citing paper
cited int64 corpus id of the cited paper
date date the citing paper's date in features_arxiv

The citation graph restricted to papers in scope: both ends are in features_arxiv, so it can be used directly as paper-to-paper structure. It covers 507,325 citing and 379,972 cited papers.

The edge is dated by the citing paper's date as features_arxiv carries it, so that a paper has one date throughout the dataset. The events tables below date citations by the citing paper's S2 publication date instead; the two agree for 99.9 percent of these edges and differ only where arXiv posted the citing paper before S2's date. Every edge has a date, since every citing paper is in scope. Self-citations of a record by itself, an S2 merge artifact, are removed.

events_citation_arxiv, events_citation_paper

Column Type Meaning
corpusid int64 the cited paper
date date, nullable the day the citations are attributed to
year int16, nullable the year, for citations known only by year
bucket enum how the citation relates to the cited paper's date (below)
citations int32 how many citations arrived on that (date, bucket)

A citation is dated by the publication date of the citing paper, since the citation edges themselves carry no date. Every citation in the S2 dump into a paper of the table is represented exactly once, in one of five buckets:

bucket date year Meaning
post set null dated on or after the cited paper's date: an ordinary citation
prepub_near set null dated up to 60 days before the cited paper's date. Usually a dating artifact: S2 defaults an unknown day to the 1st of the month, and arXiv and S2 dates differ by a median of 26 days
prepub_far set null dated more than 60 days before the cited paper's date. The paper was public elsewhere before its arXiv posting
year_only null set the citing paper carries a year but no day
null null null the citing paper has neither a date nor a year

Nothing is dropped. A consumer who wants ordinary citations filters on bucket == "post" and can still see what that excluded. A consumer who wants the year-only citations on a timeline chooses their own rule for placing them: a uniform random day within the year, the year's midpoint, or the paper's own date when the years coincide.

features_repo

Column Type Meaning
repo_id int32 surrogate key used by the other GitHub tables; stable within this release only
full_name string owner/name as on GitHub; the durable identifier
owner, name string the two halves of full_name
created date GitHub's creation date for the repository

edges_repo_arxiv

Column Type Meaning
repo_id int32 the repository
arxiv_id string the paper the README links to
date date the day the link appeared (added = true) or disappeared (added = false)
added bool direction of the change

Links are recovered from README text: arxiv.org URLs in all their forms, arXiv:NNNN.NNNNN inline references, and eprint fields of BibTeX entries. The README is read from the repository's commit history, so a link is dated by the commit that introduced or removed it, collapsed to one state per day. Three consequences:

  • The link set is a change log. The links a repository has on day D are the pairs whose latest event on or before D is an add. Roughly 30 percent of rows are removals; paper feed repositories in particular rotate their links daily.

arxiv_id is not restricted to papers in features_arxiv. Every arXiv link found is kept, in any field, so the table can be joined to any arXiv-keyed dataset. Within this release, join on features_arxiv.id to restrict to the in-scope papers.

events_stars, events_forks, events_issues, events_prs

Column Type Meaning
repo_id int32 the repository
date date the day
count int32 how many stars / forks / issues / pull requests the repository gained that day

Counts are daily gains, not running totals, and days with no gain have no row. Stars come from GitHub's per-star starredAt stream and were never decremented for unstars, so a repository's summed stars can slightly exceed its displayed count.

Quality filters applied to the paper set

These remove papers whose metadata is wrong in a way that would corrupt every date-based quantity derived from them.

  • Author-count disagreement. Papers where arXiv and S2 disagree by more than two authors are dropped: one of the two records is typically a different paper or a corrupted merge. This is not a common case.
  • Late arXiv postings. Papers with at least 25 citations, of which at least a quarter predate the paper's own date, are dropped. This catches well cited papers that were published in conferences, then uploaded to arXiv, and S2 attributes the arXiv date instead of the conference one.
  • Papers with no usable author id are dropped, since they would carry an empty author record that disagrees with n_authors.
  • Papers dated before 1991-08-14, the day arXiv opened, are dropped as impossible.

Papers excluded by the first two filters are also kept out of the author neighbourhood, so that they cannot re-enter through their co-authors.

Duplicate S2 records for one arXiv id, which happen when a preprint and its published version are never merged, are resolved to the lowest corpus id: the record created first and the one citations accumulate against.

Known limitations

  • Citation timing is reconstructed from the citing paper's date, so a citation from a paper S2 dates to a journal issue can appear months after the work was actually circulating, and S2's own dating errors propagate.
  • The crawl finds repositories through GitHub's README search, which indexes only the default branch's README
  • Link detection is textual. A repository mentioning a paper in a reading list and one implementing it produce the same edge. Repositories that link many papers can be identified from the edge table and treated separately.
  • Author identity is Semantic Scholar's author disambiguation, with its known splitting and merging errors.

Sources, licensing and availability

GitScholar is available at https://huggingface.co/datasets/huawei-csl/GitScholar and is released under the Open Data Commons Attribution License v1.0 (ODC-By), which allows reuse and modification with appropriate attribution. The full license text is in the LICENSE file alongside the data and at https://opendatacommons.org/licenses/by/1-0/. Attribution should cite the paper given in the Citation section below.

It combines:

  • arXiv paper metadata, obtained through the arXiv Open Archives Initiative (OAI) interface. Only identifiers and dates are redistributed here; titles, abstracts and categories remain in the public arXiv metadata and join on the arXiv id.
  • Semantic Scholar paper, author and citation records from the Semantic Scholar Open Research Corpus (Kinney et al., 2023), bulk release of 2026-06-16.
  • GitHub public repository metadata, collected through the official GitHub GraphQL API in accordance with GitHub's API Terms of Use. Only public metadata is included: repository names, creation dates, the arXiv identifiers found in READMEs, and daily counts of stars, forks, issues and pull requests. No personal data, no user-generated content such as README text, issue text or commit messages, and no private repository information is included.

Citation

GitScholar is described in:

GitScholar: A Dataset for Predicting AI Research Impact from GitHub Engagement. Emilien Guandalino, Lorenz K. Müller, Beatrice Alessandra Motetti, Konstantin Berestizshevsky and Lukas Cavigelli. arXiv:2609.26361, 2026. https://arxiv.org/abs/2609.26361

@article{guandalino2026gitscholar,
  title   = {GitScholar: A Dataset for Predicting AI Research Impact from GitHub Engagement},
  author  = {Guandalino, Emilien and Müller, Lorenz K. and Motetti, Beatrice Alessandra
             and Berestizshevsky, Konstantin and Cavigelli, Lukas},
  journal = {arXiv preprint arXiv:2609.26361},
  year    = {2026},
  url     = {https://arxiv.org/abs/2609.26361}
}

GitScholar is a derived database of the Semantic Scholar Academic Graph, whose ODC-By attribution requirement carries over. Please also cite:

The Semantic Scholar Open Data Platform. Rodney Kinney, Chloe Anastasiades, Russell Authur, Iz Beltagy, Jonathan Bragg, Alexandra Buraczynski, Isabel Cachola, Stefan Candra, Yoganand Chandrasekhar, Arman Cohan, et al. arXiv:2301.10140, 2023. https://arxiv.org/abs/2301.10140

@article{kinney2023semantic,
  title   = {The Semantic Scholar Open Data Platform},
  author  = {Kinney, Rodney and Anastasiades, Chloe and Authur, Russell and Beltagy, Iz and
             Bragg, Jonathan and Buraczynski, Alexandra and Cachola, Isabel and Candra, Stefan
             and Chandrasekhar, Yoganand and Cohan, Arman and others},
  journal = {arXiv preprint arXiv:2301.10140},
  year    = {2023},
  url     = {https://arxiv.org/abs/2301.10140}
}

Reading the data

All files are standard Parquet and open with any Parquet reader. In Python:

import polars as pl

papers = pl.read_parquet("features_arxiv.parquet")
events = pl.read_parquet("events_citation_arxiv.parquet")

# Citations at one year after publication, ordinary citations only
cit_1y = (
    events.filter(pl.col("bucket") == "post")
    .join(papers.select("corpusid", pl.col("date").alias("pub_date")), on="corpusid")
    .filter(pl.col("date") <= pl.col("pub_date").dt.offset_by("1y"))
    .group_by("corpusid").agg(pl.col("citations").sum())
)

# Stars of every repository linking a paper, as of a cutoff date
cutoff = pl.date(2025, 6, 17)
links = (
    pl.read_parquet("edges_repo_arxiv.parquet")
    .filter(pl.col("date") <= cutoff).sort("date")
    .group_by("repo_id", "arxiv_id").agg(pl.col("added").last())
    .filter("added")
)
stars = (
    pl.read_parquet("events_stars.parquet")
    .filter(pl.col("date") <= cutoff)
    .group_by("repo_id").agg(pl.col("count").sum().alias("stars"))
)
stars_per_paper = links.join(stars, on="repo_id", how="left").group_by("arxiv_id").agg(pl.col("stars").sum())
Downloads last month
411

Papers for huawei-csl/GitScholar