Open Source · Trends
Stack Overflow Monitoring
A data pipeline that tracks technology adoption trends by joining Stack Overflow's public data dump, the annual Developer Survey, PyPI download stats, and the dbt Hub registry.
The problem
Most "what's hot in data" lists are vendor-driven or vibes-driven. They reflect whoever has a marketing budget, not what engineers actually use or want to learn. The signal-to-noise ratio on technology trend reporting is low enough that it is nearly useless for any decision that involves real money — hiring, tooling investment, training time.
This project replaces that noise with a reproducible signal built from public, free, primary sources. The inputs are the same data everyone ignores because joining them takes an afternoon.
What it produces
The Hype Gap report
The flagship output is a ranked table of technologies sorted by the delta between "want to learn" and "used professionally" responses in the Stack Overflow Developer Survey. A positive gap means engineers want it but companies have not adopted it. A negative gap — the regret index — means it is widely used in production but nobody would choose it if they had a vote. Both directions are useful.
Signals tracked
| Signal | Source | Insight |
|---|---|---|
| Technology usage gaps | SO Developer Survey | "Used professionally" vs "Want to learn" delta |
| Community momentum | SO public data dump | Question volume, answer rates, tag velocity |
| Ecosystem growth | PyPI download stats | Real adoption vs marketing noise |
| dbt package trends | dbt Hub registry | Which data tools are actually growing |
Output formats are CSV, Markdown, and JSON. No paid APIs. No vendor data.
Data sources
- Stack Overflow Developer Survey — annual release, ~65,000 respondents. insights.stackoverflow.com/survey ↗
- Stack Overflow public data dump — quarterly release via SEDE and archive.org. archive.org/details/stackexchange ↗
- pypistats.org public API — 180-day rolling download window, no auth required. pypistats.org/api ↗
- dbt Hub registry — package metadata and download counts. hub.getdbt.com ↗
Methodology
The pipeline fetches each source independently and normalizes technology names to a
shared slug list before joining. Survey data is the anchor: every technology in the
output must appear in at least one survey year. The hype gap is computed as
(want_to_learn_pct − used_professionally_pct) per technology per survey
year, then ranked descending. PyPI and dbt counts are joined on the slug and used as
secondary signals to distinguish "genuinely growing" from "genuinely declining"
regardless of survey sentiment. Refresh cadence is annual for the survey, quarterly
for the SO dump, and on-demand for PyPI and dbt. All intermediate files are written
to disk as CSV so any step can be rerun in isolation without re-fetching upstream
sources.
Sample output
Hype Gap report, truncated. Positive gap = engineers want it, companies have not adopted it. Negative gap = widely used, low desire to continue.
Hype Gap Report — Stack Overflow Developer Survey 2024
Generated: 2024-11-03
Technology Want (%) Used (%) Gap Direction
────────────────────────────────────────────────────────────
Rust 29.8 8.4 +21.4 wanted
Kotlin 26.1 9.3 +16.8 wanted
Go 26.3 13.2 +13.1 wanted
Elixir 11.4 3.8 +7.6 wanted
TypeScript 37.1 42.5 -5.4 regret
PHP 8.9 19.6 -10.7 regret
COBOL 1.2 12.4 -11.2 regret
7 of 34 technologies shown. Full output: hype_gap_2024.csv Links
- Repository — github.com/mryagerr/stackoverflow-monitoring
- Run instructions — README usage section
- Questions or feedback — contact page
Open source, actively maintained. MIT license.