slopticgrades any live web appsign in / up

What hackathon apps look like

The results of 1,625 apps in 80 hackathons graded objectively

Almost nothing is clean

Only one app scored 0. The median is 50, and a quarter scored above 77.4. In other words, there is something wrong with almost every app.

0 to 10: 4 apps10 to 20: 199 apps20 to 30: 210 apps30 to 40: 197 apps40 to 50: 201 apps50 to 60: 169 apps60 to 70: 139 apps70 to 80: 128 apps80 to 90: 78 apps90 to 100: 74 apps100 to 110: 71 apps110 to 120: 31 apps120 to 130: 42 apps130 to 140: 18 apps140 to 150: 16 apps150 to 160: 12 apps160 to 170: 9 apps170 to 180: 11 apps180 to 190: 3 apps190 to 200: 4 apps200 to 210: 3 apps210 to 220: 3 apps220 to 230: 2 apps230 to 240: 1 appsmedian 50Q3 at 77.4050100150200slop score

More stats

What kinds of problems do apps have?

Findings have various different severities. While most are chronic, a nontrivial number of apps have serious, severe, or even critical problems. The table below shows the number of findings of each kind as well as how many apps have at least one of them.

penaltybandfindingsshareapps with at least one
1-10minor7,07675.4%1,618 (99.6%)
11-20moderate8939.5%757 (46.6%)
21-30serious6607.0%563 (34.6%)
31-40severe2672.8%248 (15.3%)
41+critical4855.2%420 (25.8%)
all9,381100%(these overlap)

Grading an app by its single worst finding gives the stats below. Each one contains the ones under it, so almost 3 in 5 projects carry a significant problem and virtually every app overlooks some hygiene.

(A significant problem has a penalty more than 20, and an acute problem has a penalty of more than 40.)

Winners ship more slop

Counterintuitively, winning apps have 11.8% higher median slop than the rest.

54.9median slop, winners253 apps
49.1median slop, everyone else1,372 apps

The same is true for Lighthouse:

74median Lighthouse score, winners253 apps
80median Lighthouse score, everyone else1,372 apps

As you can see, winning does not correlate with app cleanliness. In fact, the opposite tends to be true. Most hackathons employ human judging, which rewards ideas, features, presentation, and the demo over durability. Winning apps tend to ship more features, meaning more surfaces to misconfigure or get wrong, and human judges do not have time to judge quality consistently over hundreds of apps.

Fast != clean

Lighthouse performance barely predicts anything else. Measured against slop with the performance axis taken out, the correlation is -0.071 across 1,571 apps, which is close enough to zero to call the two independent. In other words, speed and durability don't have any relationship.

Breakdown per hackathon

Across 61 hackathons out of 80 with 8 or more graded apps, median slop runs from 28.3 to 105.8, a 3.7x difference. Hover over a bar for the event in question.

each row is one event, sorted by median

Yet exploits are rare

Only 2.9% of apps carry something an attacker could use today. The largest single finding is an exposed backend, with 18 apps serving a Supabase or Firebase database that anyone could read, because row level security was not turned on. Of those, 7 returned records in bulk and 3 held personal data in its columns.

what this needs

Note that to find an exposed backend or other exploitable vulnerability, Sloptic must grade actively. Grading passively holds back active attacks, at the cost of missing exploitable findings. To grade actively, verify your domain or event.

What didn't get graded

Sloptic attempted 2,685 apps and graded 1,625 of them, or 60.5%. Most of the rest were due to link rot (expired free tier), timeouts, a WAF challenge, or other reasons.

why an app was not gradedapps
dead URL (link rot / 4xx / 5xx)757
ungraded (grade aborted / timed out)113
other100
entry challenge (WAF withheld the grade)90
not graded1,060

Also excluded are 82 streamlit apps since Sloptic is currently unable to properly separate what the teams built from these platforms.

How to read this

Every figure here is an aggregate. No apps were named to protect the privacy of individual teams that built them. The apps in this study were graded as a calibration to build Sloptic itself.

These are the full grade numbers that comprise the corpus used for percentile ranking on active grades. A separate curve exists for passive grading.