What hackathon apps look like
The results of 1,625 apps in 80 hackathons graded objectively
Almost nothing is clean
Only one app scored 0. The median is 50, and a quarter scored above 77.4. In other words, there is something wrong with almost every app.
More stats
- 58.5average slop.
- 38standard deviation.
- 239.1the worst app.
What kinds of problems do apps have?
Findings have various different severities. While most are chronic, a nontrivial number of apps have serious, severe, or even critical problems. The table below shows the number of findings of each kind as well as how many apps have at least one of them.
| penalty | band | findings | share | apps with at least one | |
|---|---|---|---|---|---|
| 1-10 | minor | 7,076 | 75.4% | 1,618 (99.6%) | |
| 11-20 | moderate | 893 | 9.5% | 757 (46.6%) | |
| 21-30 | serious | 660 | 7.0% | 563 (34.6%) | |
| 31-40 | severe | 267 | 2.8% | 248 (15.3%) | |
| 41+ | critical | 485 | 5.2% | 420 (25.8%) | |
| all | 9,381 | 100% | (these overlap) | ||
Grading an app by its single worst finding gives the stats below. Each one contains the ones under it, so almost 3 in 5 projects carry a significant problem and virtually every app overlooks some hygiene.
(A significant problem has a penalty more than 20, and an acute problem has a penalty of more than 40.)
- Hygiene99.9%1,624 apps
E.g. missing headers, poor accessibility, orange Lighthouse score; the things people often skip.
- Significant58.9%957 apps
Findings that are not just cosmetic, such as dead controls, broken links, missing rate limits, overly slow pages.
- Acute25.8%420 apps
Severe issues that noticeably degrade the user experience or allow attacker access, such as crashes, unusable pages, or exposed backends.
- Exploitable2.9%47 apps
Catastrophic vulnerabilities that an attacker can exploit today.
Winners ship more slop
Counterintuitively, winning apps have 11.8% higher median slop than the rest.
The same is true for Lighthouse:
As you can see, winning does not correlate with app cleanliness. In fact, the opposite tends to be true. Most hackathons employ human judging, which rewards ideas, features, presentation, and the demo over durability. Winning apps tend to ship more features, meaning more surfaces to misconfigure or get wrong, and human judges do not have time to judge quality consistently over hundreds of apps.
Fast != clean
Lighthouse performance barely predicts anything else. Measured against slop with the performance axis taken out, the correlation is -0.071 across 1,571 apps, which is close enough to zero to call the two independent. In other words, speed and durability don't have any relationship.
- -0.071Spearman correlation between Lighthouse performance and the rest of the slop (without the performance component).
Breakdown per hackathon
Across 61 hackathons out of 80 with 8 or more graded apps, median slop runs from 28.3 to 105.8, a 3.7x difference. Hover over a bar for the event in question.
Yet exploits are rare
Only 2.9% of apps carry something an attacker could use today. The largest single finding is an exposed backend, with 18 apps serving a Supabase or Firebase database that anyone could read, because row level security was not turned on. Of those, 7 returned records in bulk and 3 held personal data in its columns.
what this needs
Note that to find an exposed backend or other exploitable vulnerability, Sloptic must grade actively. Grading passively holds back active attacks, at the cost of missing exploitable findings. To grade actively, verify your domain or event.
What didn't get graded
Sloptic attempted 2,685 apps and graded 1,625 of them, or 60.5%. Most of the rest were due to link rot (expired free tier), timeouts, a WAF challenge, or other reasons.
| why an app was not graded | apps |
|---|---|
| dead URL (link rot / 4xx / 5xx) | 757 |
| ungraded (grade aborted / timed out) | 113 |
| other | 100 |
| entry challenge (WAF withheld the grade) | 90 |
| not graded | 1,060 |
Also excluded are 82 streamlit apps since Sloptic is currently unable to properly separate what the teams built from these platforms.
How to read this
Every figure here is an aggregate. No apps were named to protect the privacy of individual teams that built them. The apps in this study were graded as a calibration to build Sloptic itself.
These are the full grade numbers that comprise the corpus used for percentile ranking on active grades. A separate curve exists for passive grading.