Skip to content

Ilia DudaCo-op Jan 2027

Ranking startup segments, and how much the answer depends on the data

Data-analytics capstone, Yandex Practicum · December 2025 · Python, pandas

A composite model ranks 48 startup segments on growth and size across 35,589 funding records. Re-executing my capstone exactly, then changing one data decision at a time, measures how much the recommendation depends on them — three treatments give three different top picks — and finds the segments that hold up under all three.


Fig. 1
One model, three treatments of the data, three top pickscomposite score 0–100 · 48 mass segments · growth 0.35 · CAGR 0.30 · funding 0.20 · companies 0.15

Segments missing from the growth table score zero on 65% of the weight.

indigo, and set bold: in the top fifteen under all three treatments · ∅ absent from the notebook’s growth table, so scored zero on growth and CAGR · arrows: places moved against as written

  1. 1Technology56.9
  2. 2Apps39.0
  3. 3Medical36.8
  4. 4Startups33.1
  5. 5SaaS32.6
  6. 6Biotechnology∅31.3
  7. 7Design29.7
  8. 8Big Data27.8
  9. 9Software∅24.7
  10. 10Internet17.7
  11. 11Manufacturing12.1
  12. 12Health Care∅11.3
  13. 13Real Estate11.3
  14. 14Clean Technology∅10.9
  15. 15Mobile∅10.8

Top segments by composite score under three treatments of the same 35,589 funding records. As written, the top three are Technology, Apps, Medical. With the growth lookup fixed the top pick becomes Clean Technology, and the rank correlation with the original falls to 0.18. With the CAGR start fixed as well, the top pick is Software.

Treatment
As written
Top pick
Technology
Spearman ρ against as written
1.00
Scored zero on growth
38 of 48
The weights never change; only the handling of the data does. Fixing the growth lookup alone leaves the ranking almost uncorrelated with the capstone’s original (Spearman ρ = 0.18), which is the whole argument for treating a model’s data pipeline as part of the model.
Rank of each segment under the three treatments
SegmentAs writtenGrowth lookup fixedLookup and CAGR fixed
Technology1433
Apps2115
Medical33840
Startups42013
SaaS51016
Biotechnology6323
Design71330
Big Data81215
Software9181
Internet103934
Manufacturing113227
Health Care124628
Real Estate133618
Clean Technology14112
Mobile152811
Enterprise Software164224
E-Commerce172919
Curated Web18276
Semiconductors19236
Advertising203425
Hardware + Software21617
Games223122
Finance233721
Health and Wellness243520
Web Hosting254847
Security264037
Analytics274744
Social Media28710
Education2952
Hospitality3083
Fashion31929
Messaging322346
Consulting334335
Travel34304
News351414
Search364132
Music372242
Video381731
Photography391938
Automotive401626
Sports41219
Cloud Computing424443
Networking431545
Marketplaces44248
Public Relations452648
Entertainment46337
Social Network Media472539
Nonprofits484541

The question

Which startup market segments should an investor favour for the coming year? The analysis starts from 54,294 funding records spanning 2000 to 2014 and ends at a recommendation, and every step in between is a decision about data: what counts as a duplicate, what to do with a missing date, which extreme rounds to trust, how to measure growth, and how to weigh it against sheer size.

Fig. 2
From supplied records to scored recordsstartup-investment-analysis · each stage of cleaning, and what it removed
From supplied records to scored recordsOf 54,294 supplied records, 40,907 carry both a funding record and a date. Removing outliers by per-segment interquartile fences and dropping years with too few funding rounds leaves 35,589 records in the ranking — 34.5 percent of the original set is excluded.records as supplied 54,294− 13,387 lacking a funding record or a datewith funding and dates 40,907− 5,244 outliers, − 74 in thin yearsscored 35,58934.5% of supplied records are not ranked
Duplicates and rounds without funding go first; then 5,244 outliers by per-segment interquartile fences, so a segment of seed rounds is not judged by the fences of a segment of growth rounds; then years with too few rounds to measure.

The model

Segments are sized by company count into mass, mid-size and niche; the ranking covers the 48 mass segments of 395. Each is scored on four components — funding growth into 2014, compound annual growth, total funding and company count — each min-max scaled to 0–100 across the segments and combined with weights of 0.35, 0.3, 0.2 and 0.15. As written, the model recommends Technology, Apps and Medical.

Stress-testing it

A recommendation is only as good as its sensitivity to the choices underneath it, so I re-ran my capstone exactly and then changed one data-handling decision at a time. The growth lookup, as written, only knows the 10 segments whose funding rose in 2014, so the other 38 score zero on 65% of the weight — including Biotechnology and Software, the two largest segments by funding. Giving every segment its real growth, negative where funding fell, reorders the ranking almost completely.

The compound growth has a second sensitivity: measured from 2000, a segment with no funding that year starts from a dollar, which inflates its growth rate into the hundreds of per cent. Measuring from each segment’s first funded year moves the top pick again, to Software. And 2014 itself is incomplete in the source data — its last two months are a fraction of a normal month — so part of every “decline” is under-reporting rather than a market signal.

What survives

No segment is in the top ten under all three treatments. Only Apps, Clean Technology and Big Data are in the top fifteen under every one of them, which makes them the defensible recommendation: the answer that does not depend on which of three reasonable data decisions you believe. A model’s data pipeline is part of the model, and a ranking is only as useful as the report of what it is sensitive to.


data
the course’s startup investment dataset, 54,294 records, hash-pinned in the re-run
stack
Python · pandas · NumPy · matplotlib · Jupyter