How the values are built
Every number on this site is arithmetic you could run yourself. Same data in, same number out, every time — no language model is anywhere near it. This page is the whole recipe: which sources outrank which, what the estimator is, exactly what it is fed, and the things it is provably bad at.
First
Four rungs, in this order
For every graduate we walk down this ladder and stop at the first rung that holds. The badge next to his value tells you which one it was, on every single card — a number without its provenance is not published here.
- 1
A reported valuation fact, with its receipt
REPORTEDSomeone bid €X and was turned down. A contract announcement names a release clause. A club states a price. These are facts about the market, reported by a named outlet on a dated page, and they beat everything else — because they are the only evidence that exists for players nobody has actually sold. Each one is entered by hand with its source URL and an expiry date, and the card on the site prints the source. When the expiry passes, the anchor is deleted rather than extended, and the player drops back to the model.
- 2
A real transfer fee, actually paid
FEEIf a club paid a disclosed fee for him in the last twelve months, that fee is the value. Full stop, no modelling. Between twelve and thirty-six months the fee stops being the answer and becomes the floor: a nineteen-year-old who has exploded since his move may be worth more than his price tag, but nobody is ever valued below what someone recently paid. Past thirty-six months a fee is history, and the card says so.
- 3
The model's estimate
ESTNo fee, no reported fact — so we estimate. That estimate comes from decision trees trained on fees that were really paid and really disclosed, described in full below. This is where most of the site's numbers come from, and it is the only rung where the number is ours rather than someone else's.
- 4
A published formula, as the last resort
ESTWhen a player's record is too thin for the model — no date of birth, no resolvable position — a plain formula steps in: a base value from an age table, multiplied by a position factor, multiplied by his club's strength, boosted by his international caps. Every constant is in a public config file. It is crude on purpose, and it covers about one per cent of graduates.
Rungs 1 and 2 are somebody else’s number and we print their receipt. Rungs 3 and 4 are ours, and they wear the EST badge so you always know the difference.
Rung three
What the model actually is
It is a gradient-boosted decision tree ensemble — a few hundred small trees, each one correcting the last one’s mistakes. It is not a neural network, it does not read text, and it has no opinions. It has only ever been shown one kind of thing: transfer fees that were really paid and publicly disclosed, parsed from Wikipedia’s per-window transfer lists, together with the buying club, the date, and the player’s record.
Old fees are restated in today’s money first. The correction is not an economic index we found somewhere — it is derived from our own transfer table, as the median disclosed fee per twelve-month window, smoothed, with the most recent window set to 1.0. Football inflation measured with football fees.
The model predicts the logarithm of the fee, not the fee, so it is fitted on proportional error: being €4m out on a €10m player is the same size of mistake as being €40m out on a €100m one. The result is then clamped between €0.5m and €222m, the all-time record fee — a guardrail, not a judgement. One editorial constant sits on the end: a market-level multiplier, currently 1, which exists so the whole output can be re-levelled against published press aggregates if it drifts. At 1 it does nothing at all.
388
fees trained on
2023-03-08 → 2026-07-30
17
inputs per player
listed in full below
€18.4m
mean error
on 50 real fees held back from training
€10.2m
median error
half our estimates land closer than this
€146.3m
biggest fee it has seen
it cannot predict above this
9
fees over €80m
in the entire corpus
The inputs
All 17 of them, and nothing else
This is the complete list. There is no scout report in here, no hidden adjustment, and no field where anybody types a number they like. If a fact about a player is not on this list, it did not move his value.
- Age
age - His age on the valuation date, to the decimal. The strongest single signal in football pricing: the same player is a different asset at 22 and at 31.
- Position
pos_GK · pos_DEF · pos_MID · pos_FWD - Four on/off switches — goalkeeper, defender, midfielder, forward — of which exactly one is lit. Keepers and forwards have never been priced alike and the model is not asked to pretend otherwise.
- Career appearances
total_apps - Every senior and loan appearance across his career, added up. Proof he plays, not just that he is signed.
- Career goals
total_goals - Goals in those appearances. Worth a great deal for a striker and very little for a centre-back, which is why the model sees position too.
- Appearances at a club on this board
ranked_apps - How many of those games were for the FIRST team of one of our 160 academies. Reserve sides — Barcelona B, Bayern II, Real Madrid Castilla — are stripped out, because fourth-tier reserve football counting as "played at a big club" was quietly inflating a fifth of this feature.
- Appearances per year since his debut
apps_per_year - Playing time as a rate rather than a total, so a 33-year-old squad filler does not out-score a 21-year-old who starts every week.
- Appearances per year of adulthood
age_apps - Appearances divided by (age − 16): how much senior football he has played for how young he is. The signature of a player brought through early.
- Senior international caps
senior_capsone-way - Games for the full national team. The clearest public verdict on a player that does not come from a transfer market.
- Youth international caps
youth_caps - U-17 to U-21 selections. What his country thought of him before anyone had bought him — which is most of what there is to know about a teenager.
- Honours won
honours_countone-way - How many honours his Wikidata record lists.
- Honours, weighted
honours_weightone-way - The same honours, but a Ballon d'Or is not a league title. The weights are in a public config file: Ballon d'Or 30, a World Cup 15, the Kopa Trophy 15, a Champions League 8, a domestic title 3, anything unlisted 1.
- Wikipedia readership
log_pageviewsone-way - The logarithm of his median monthly Wikipedia pageviews. Our only measure of public attention — how many people go and look him up — and a decent stand-in for the noise around a player that fee data alone never captures.
- Famous and young
fame_youthone-way - Readership multiplied by (24 − age), floored at zero. It exists because the fee history contains no sale of a generational teenager at all, so nothing else in the model can reach those numbers. It is the one input that is openly an extrapolation — and the reason the ladder above puts reported facts on the top rung.
- His current club's strength
elo - The Elo rating of the club he plays for today, from ClubElo, falling back to a published per-league table. If neither can place his club, the model is handed a BLANK rather than a low number — a club we cannot identify is missing information, not a bad club, and conflating the two was worth a systematic error.
The inputs marked one-way carry a hard constraint: the model is forbidden from ever lowering a value when one of them goes up. Winning another cap, another trophy, or gaining readers cannot make a player cheaper. That is not something the data taught it — it is a rule we imposed, because the alternative is a model that occasionally produces an insulting number for a reason no one can explain.
The uncomfortable part
What it cannot do
It cannot price a player nobody has sold. Decision trees do not extrapolate. Whatever the largest fee in their training data is, that is the highest number they can produce, for anyone, ever. Ours is €146.3m, and only 9 disclosed fees in the whole corpus clear €80m. So the very best players — the ones whose clubs would never sell them, who therefore generate no fee for anyone to learn from — come out systematically low. This is not a bug we are about to fix; it is what a fee-trained estimator is. It is also the entire reason rung 1 exists: when the press reports a rejected bid or a release clause, we use that instead and show you the article.
It is often wrong, and we publish by how much. Every run trains on the older fees, then predicts the last 120 days of real disclosed fees it has never seen — 50 of them. The average miss is €18.4m and the median miss €10.2m, the gap between them being the handful of big transfers it gets badly wrong. Treat every EST on this site as a number with that much air around it.
Its view of the market is English-language. Fees are parsed from Wikipedia’s “List of English football transfers” pages, because those are the ones that carry machine- readable amounts. Deals between two Serie A clubs, or two Brazilian ones, are largely invisible to it — so it has learned what the Premier League pays more thoroughly than what anyone else pays.
It reads today’s record on yesterday’s transfer. Appearances, goals, honours and readership are current totals, used even for a fee paid three years ago — no historical snapshot of those numbers exists in open data. Club strength is likewise today’s Elo table, not the table on the day of the deal. Both make the model slightly too clever about the past, and we would rather say so than bury it.
It has never seen a fee that was not disclosed. Undisclosed fees, add-ons, sell-on clauses and swap deals are all outside its world. The number it produces is an estimate of a headline fee, not of a club’s accounting.
When we get one badly wrong and the player then moves, the miss goes in the ledger with the real fee next to it. The track record keeps the bad calls as carefully as the good ones.
You asked directly
No language model touches the number
The valuation pipeline is Python, SQL and scikit-learn. There is no model of language in it, no API call to one, and no place where free text becomes a euro amount. Run it twice on the same snapshot and you get the same numbers to the cent: the training data is fixed by the snapshot date, the algorithm is deterministic, and the one random element — the seed the trees are built with — is pinned to a constant in the source. The estimate on a card is a pure function of the published data.
A language model does appear in this project, once, in a role with no write access to anything: as an auditor. It reads a published value, searches the web, and reports whether a named outlet has published a relevant FACT — a bid, a clause, a club statement, a fee. It returns flags and URLs. It is explicitly forbidden from supplying a valuation of its own, and if the only figure it can find traces back to a valuation aggregator, the required answer is “no fact found” and our number stands unchanged. Anything it flags is read by a human, who writes the anchor by hand with the source URL and an expiry date, or does not.
An LLM that “recalls” a player’s market value is quoting an aggregator from its training data. That is why it is never asked. No value from Transfermarkt or any other valuation database is scraped, stored, cited or laundered through anything in this project; the whole dataset is built from Wikidata (CC0), Wikipedia (CC BY-SA), ClubElo and named press articles.
You do not have to take any of this on faith — the code is public: the model, the valuation ladder, every anchor with its source and the auditor’s rules.
If we are wrong
How to challenge a number
Open an issue. Tell us the player and what you think is wrong, and if you have one, attach the thing that settles it: a link to an article reporting a bid, a clause, a completed transfer or a club statement, with its date. That is what we can act on — a reported fact becomes an anchor at the next refresh, and the card starts showing your source.
What we cannot act on is a figure from a valuation site, however confident. It is not a fact about the market, and it is not something we are able to cite. “He is obviously worth more than that” is welcome too, mind — it is often right, and if it is right about several players at once it is a model problem rather than a data problem, which is more useful to us than any single correction.
Open an issue · corrections ship with the next refresh, each one recorded in a public file with its reason. Wrong positions, missing graduates and eligibility arguments go to the same place — the rules cover those.
Data as of 2026-08-03 · rules phase0-r1r16-v1 · values gbt-v2 · model fitted 2026-08-03. Not affiliated with any club, league or federation.