Scouting Uncertainty: adding value to data scouting results

9 min read Original article ↗

Marc Lamberts

Press enter or click to view image in full size

I have been looking at spreadsheets for football scouting for over a decade, and every now and then, my perception completely changes. The biggest change I have experienced is when I started to realise that we use data mostly to shape profiles and how they fit in the profile. Pure quality can be measured, but it often is so conflicted with different variables, that it is difficult to make a definitive contribution.

Lately, I’ve been creating more and more on a meta level, mostly in data engineering. In that light, I wanted to share something I have introduced in my day-to-day data scouting when working with aggregated data: scouting uncertainty.

The idea is to add an uncertainty score to the scouting score to add a level of trust towards the data. If the data is trustworthy and uncertainty is low, the quality of data scouting for that specific player will be higher.

Data

For this little research, I have had a look at Wyscout’s aggregated data of the J2-J3 league of 2026. Normally these are split into two different leagues, but for this anniversary year, things are a little bit different.

This data focuses only on strikers, so I’ve only looked at players with CF as a position and/or players with multiple positions, but CF as a dominant position in the games they have played.

The data has been collected in XLSX files from the Wyscout platform on June 15, 2026. Things will not have changed since the World Cup has been going on and the new season has to start, but for clarity's sake, here it is.

Profiles

Not every striker is created equal. What I mean by this is that a striker is the position, but the profile is different. In other words, what kind of striker is the player?

I have categorised the strikers into five distinct profiles:

  • Goalscoring Striker
  • Target Striker
  • Dynamic Striker
  • Second Striker
  • Complete forward

All these profiles are created from all the data we have available, but are created from weighted composite scores in z-scores, which are normalised from 0–100. This doesn’t say anything about the quality of the striker, but much more about the tendencies of said striker. A score of 67 means that this player scores 67% on the perfect striker profile scale.

Press enter or click to view image in full size

We have 70 players that fit the filters, and if we look at the image above, we can state the following. The Target Striker leads the profiles with 31%, followed by the Goalscoring Striker with 23%, and the Second Striker with 23%. The Dynamic Striker has 20%, and finally, we end with a complete forward with only 3%. Note for the complete forward, they are only complete forwards if they score 65% or higher in each of the categories, which is a strict cutoff.

Press enter or click to view image in full size

In the table above, you can see all 70 players and how they rank in the particular categories of striker profiles. This is more or less just information about each player and nothing definitive in terms of quality. What are the biggest relations are within two profiles, that is, between the Goalscoring Striker and the Target Striker?

Press enter or click to view image in full size

In the visual above, you can see the five player per category who score the highest in their respective categories. Sagawa is the best fit for Target Striker, Mori for Dynamic Striker, Nishino for Goalscoring Striker and Anderson for Second Striker.

Calculating the uncertainty score

Every Scouting Score in this model comes paired with an Uncertainty Score, a 0 to 100 rating of how much to trust the number, where higher means less reliable. The idea is simple: a striker who has racked up 90 in a full, injury-free season means something very different from a striker who has hit the same number in half as many minutes off a hot streak of shooting luck. Rather than quietly averaging over that risk, the model surfaces it as its own figure, so a big Scouting Score can be read alongside a clear-eyed sense of how much of it is signal versus noise.

Press enter or click to view image in full size

The Uncertainty Score is built from five factors, each scored 0 to 100 and blended into a weighted total:

  • how much game time the player has actually logged relative to a settled 20-match sample (35% of the total, the single biggest driver);
  • how many shots they have taken, since shooting and conversion numbers are notoriously unstable on small volumes (20%);
  • how far outside a 24 to 29 prime age band they sit, since younger and older players’ underlying output tends to swing more from season to season (15%);
  • how much of the underlying data was simply missing (15%);
  • and whether their profile is genuinely well-rounded or propped up by one or two standout numbers rather than consistent performance across the board (15%). The result is a single confidence read, High, Medium, or Low, that sits next to every player’s Scouting Score, so a standout number always comes with an honest answer to the question of whether it can be trusted.

Press enter or click to view image in full size

If we analyse our dataset, we can find that the mean uncertainty is 30,4 on a scale of 0–100. The median is almost the same, with 30, which is a pretty solid uncertainty score. If we look at the highest category, high-confidence players, who are the best in terms of the lowest uncertainty.

Press enter or click to view image in full size

In the scatterplot above, you can see the scouting score compared to the uncertainty score. A high scouting score signifies the quality of the player according to the data, while a high uncertainty score signifies that we cannot fully trust the data. A lower score for uncertainty is better in this regard.

You can see that here as well, a high scouting score + a low uncertain score, makes for the best scouting options if we are looking purely at the data.

Press enter or click to view image in full size

Okay, but what drives that uncertainty? We have looked at the calculations, but what drives it in our league, we are researching?

  1. Minutes drive the most with 15.3 points on average. This is where we look at the minutes played throughout the season in comparison with the total.
  2. Metric spread is followed by 7.6 points on average. This is where we look at the spread of quality throughout several metrics. If a player has an even distribution of metrics, that is better than performing exceptionally well in one and weak in the rest.
  3. Shot-volume gap comes next with 5.8 points. This is where we look at the volume of shots for strikers. Yes, they can have a high shots per 90, but what’s the total volume of shots throughout the season? We want to limit the effect of an anomaly in a high-shots game.
  4. Age volatility with 1.7 points. This is where we look at peak ages and give different meanings to very young or relatively old players.
  5. Missing data. This is where we look at the data that is missing. Now here there is nothing, because the season has ended and nothing is missing. If you do this analysis during the season, this might be different due.

Results

Press enter or click to view image in full size

In the visual above, you can see the final results for all strikers in the J2-J3 league 2026. Every player’s score shows the total uncertainty score, but also how that score is built up. This shows you what component drives the score.

If we look at the “worst” players in this regard, they are driven by multiple categories, and that shows that uncertainty is higher. When we look at the “best” players, we see that they have fewer metrics. What’s most significant is that age volatility hardly plays a role.

In the two tables above, you can see the two categories of Goalscoring Striker and Dynamic Striker in their profile. The scouting score is the fit for that particular profile, and the uncertainty score gives us an idea of how much to trust that data. In other words, how much do we trust the profile in relation to the individual data of the player?

For example, in the Goalscoring Striker, the best fit is also a player whose data we can trust. But in the Dynamic Striker profile, we could go with the second one on the list, because the uncertainty is lower.

Here we see the same kind of tables, but for other profiles. The thing I think is interesting is Patric from Zweigen Kanazawa. He scores relatively high on the Target Striker profile, but his uncertainty is really high compared to the rest. That would be a player you don’t trust in terms of data.

Final thoughts

What this exercise really produces is not a ranked list of “the best striker in Japan II and III,” it is a structured way of asking two questions at once: what kind of striker is this, and how much should we trust the number attached to his name. Seventy CF-listed players sorted cleanly into recognisable profiles, Target Strikers, Goalscoring Strikers, Second Strikers, Dynamic Strikers, and a rare pair of Complete Forwards who graded well across every dimension at once. That alone is useful for a recruitment department trying to fill a specific tactical need rather than simply chasing the highest overall score.

The Uncertainty Score is the part worth dwelling on, because it resists the instinct to treat a single number as gospel. A player putting up an elite Scouting Score on 700 minutes and a handful of shots is a different proposition than one doing it across a full season, even if the two numbers look identical on a spreadsheet. The efficient frontier on the risk map makes that trade-off visible: T. Nishino, D. Furukawa, and H. Iwabuchi all sit on it, meaning nobody in the pool beats them on both quality and reliability at once, while a name like Patric shows how a modest sample and an inconsistent underlying profile can inflate uncertainty even when the raw output looks tempting.

None of this replaces a scout with a laptop and a match to watch. It is a screening layer, a way to narrow seventy players down to the ten or so worth a longer look per role, and to walk into that look already knowing whether the data behind a name is thick or thin. Treat the shortlists as a starting point for video and in-person work, not a substitute for it, and treat a high Uncertainty Score not as a red flag to discard a player but as an honest instruction to watch more of him before deciding anything.