Part 08

The decade the record missed

One university’s computing theses, 2016–2025, rebuilt from its own open repository after the commercial record went dark. Dissertations built on large language models go from nothing to a quarter of the doctoral cohort, entirely inside the missing years.

2,788 theses · doctoral and master’s · 2016–2025 · two sources merged

Part 05 measured this university against every other institution in the corpus and stopped at 2022, because from 2023 the university had almost entirely ceased depositing dissertations with the commercial database this project is built on. The university runs its own open-access repository, and it has the missing theses.

This page is the decade 2016–2025 rebuilt from both sources together. Inside that window the subscription database holds 615 doctoral theses. The merged record holds 1,377, and adds 1,411 master’s theses where there had been none: 2,788 in all, 4.5 times the material. Of the doctoral theses, 420 appear in both sources, 762 only in the repository and 195 only in the subscription database, so neither source on its own is the record.

0 40.1 80.2 120 160 200 2016 2017 2018 2019 2020 2021 2022 2023 2024 2025 Merged record, master's Merged record, doctoral Subscription database only
Computing theses at one university, 2016–2025. The lower line is what the subscription database holds; the upper is what the merged record holds. Master’s theses are absent from the subscription source entirely. The repository is missing May 2016, so the first master’s point is short by about two thirds.

The gap was never a cliff. Matching every repository record by title against the national corpus, the share of this department’s computer science dissertations that the national database also holds runs 98.8% for the 2011–2015 cohorts, 49.7% for 2016–2022 and 9.5% for 2023–2025. The 2023 outage is the same curve reaching bottom, and the years everyone has been treating as clean were already missing half the department.

The thing that was hiding in the missing years

The share of dissertations building on large language models was 1.1% or lower in every cohort through 2022, which is what you would expect since the models barely existed. Then 10.3% in 2023, 12.2% in 2024, and 24.2% in 2025. Nearly a quarter of this university’s computing doctorates now write dissertations that use them.

The three years carrying the entire signal are exactly the three the subscription record was missing. Run this analysis on the older data and the line is flat at zero, and nothing about a flat line looks broken.

0% 15.5% 31.1% 46.6% 62.2% 77.7% 2016 2017 2018 2019 2020 2021 2022 2023 2024 2025 Doctoral, machine learning Doctoral, deep learning Master's, large language models Doctoral, large language models
Share of dissertations whose abstract indicates the work uses machine learning, deep learning, or large language models specifically. Labels assigned by an open-weight model reading each abstract.

The master’s line moves earlier and further: 25.7% by 2024 against 12.2% for doctorates, and 25.8% in 2025. That ordering is what the difference in project length predicts. A master’s thesis begun in 2023 is finished in 2024; a dissertation begun in 2023 is finished around 2027, and the 2025 doctoral cohort is mostly people who chose their topic before any of this existed. On that reading the doctoral line has not peaked, and the 2025 figure is a floor rather than a level.

Underneath it the older wave is still running and has not been displaced. Deep learning went from 10.8% of dissertations in 2016 to a peak of 57.1% in 2024, and 49.2% in 2025. The language-model share is being added on top of a field that had already largely converted, not taking its place.

Does working on it change where you go?

Pooling the 2023–2025 cohorts, doctorates whose dissertations use large language models go into industry at 61.4%, those using deep learning without them at 54.9%, and those using neither at 32.8%. A tidy ranking, and largely an artefact of who graduated when: 59% of the language-model group graduated in 2025 against 35% of the neither group. Standardising all three to the same year mix:

Uses large language models 57% Uses deep learning, not LLMs 56.4% Neither 32.6%
Share entering industry, doctoral cohorts 2023–2025, standardised to a common year mix. Groups are defined by what the dissertation does, not by subject area.

The gap between the two machine-learning groups disappears, 57.0% against 56.4%, while the gap to everyone else does not. The academic side says the same thing from the other direction: 12.0% and 12.0% respectively against 30.4% for the rest. Whatever is sorting these people, it is not large language models specifically. It is whether the dissertation is a machine-learning dissertation at all, and that divide long predates them.

Nationally, on the same measurement, the three groups run 53.5%, 44.5% and 38.3%, and there the language-model group does sit clearly above the deep-learning group. With 44 people in the Illinois language-model group the national ordering is the one to trust.

Where the decade’s doctorates went

Against the national corpus measured the same way, this institution sends more of its doctorates into industry, 51.6% against 47.6%, and slightly fewer into academia, 22.8% against 24.1%. The share found working at all is 86.5% against 84.9%. The institution sits close to the national pattern, tilted a little further toward industry.

Dissertation area Industry Academia Research institute Never observed
Data and retrieval 72%
Languages and software engineering 61%
Systems 59% 23% 9%
AI and machine learning 56% 18% 4% 14%
HCI and graphics 47% 34%
Applied and scientific 44% 26% 18%
Security 41% 25%
Theory 41% 29%

Read down the industry column and the areas run from 41% to 72%; down the academic column, from 18% to 34%. That is a thirty-point spread sitting behind a choice most students make on interest alone.

Areas group the dissertation subject labels. By share of dissertations the largest are AI and machine learning at about 34%, systems about 20%, and applied and scientific computing about 10%. Each row is a percentage of the doctorates in that area who could be matched to an employment record, and that base differs by area, so rows are comparable with each other. Rows do not sum to 100: government and unclassifiable employers are not shown.

One caution about job titles. “Research” is the largest role category for doctorates here, 232 of them in industry and 126 in universities, and those are not the same job. The university group is largely postdoctoral, and the estimated pay attached to the two differs by roughly a factor of three.

Master’s theses, for the first time

The subscription source holds no master’s theses from this university at all, so this is the first time they can be looked at. 60.0% go into industry against 52.2% nationally, and 14.3% into academia. The tilt toward industry is stronger than at doctoral level, which is the expected direction.

Only students who write a thesis deposit one, and the university’s largest computing master’s programme is coursework-based, so this is the research end of the cohort rather than the whole of it. 155 of them (17%) went straight on to a doctorate at the same university and are excluded from every destination figure here, because their first year after the master’s is another year as a student.

The ten-year curve

Follow the cohorts forward and the academic share appears to trace a U: 28.5% in the year of the degree, down to 20.8% by year 3, back to 29.0% by year 10. It is a satisfying shape, and it is the shape people use when they describe the academic career to each other. A postdoctoral trough, then a return to faculty positions.

It is not real. Each point on that curve is computed from whichever cohorts have reached that age, so the right-hand end is not the same people as the left. The dark line below holds the group fixed, the same 194 people measured at every horizon, nobody entering or leaving.

0% 7.5% 15.0% 22.5% 30.0% 37.5% 0 1 2 3 4 5 6 7 8 9 10 Academia, every cohort available Academia, same 194 people throughout
Share of doctorates in academia, by years since the degree. The pale line lets the group change as cohorts age in and out; the dark line follows one fixed group of 194 people the whole way.

Held fixed, the share falls from 33.5% to 26.3% by year 2, and from there stays between 25.3% and 26.3% all the way to year 8. No trough, no return. The U was composition: earlier cohorts were more academic than later ones, so as the long horizons come to contain only earlier cohorts the average drifts upward on its own.

The national data points the same way, with the sample size to be sure of it. Split into cohort groups and followed within group, every group that has run its full course declines from beginning to end and never turns back up: 2000–2009 from 27.7% to 24.6% by year 10, 2010–2014 from 24.1% to 19.7%, 2015–2019 from 24.1% to 19.3%. The recovery exists only in the pooled series, and only because of who is left in it.

The three years the commercial record missed are the three years that changed the field. Language-model dissertations went from nothing to a quarter of the doctoral cohort, and arrived twice as fast at master’s level, on top of a deep-learning wave that had already taken half the department. And the academic career, followed on a fixed group of people, has no second act: the share in academia falls for two years after the degree and then stays flat for the next six.