The sheet arrives about twenty minutes after the whistle. One page, two columns, thirty or so rows. Possession 58 to 42, passes 512 to 341, accuracy 84 per cent against 71, shots 14 to 9, on target 5 to 3, duels won 51 to 49. Somebody screenshots it, somebody else builds an argument on it, and by Sunday the argument has hardened into a description of the match.
Most of those numbers are real. Almost none of them mean what the argument assumes. Roughly half the rows on a standard sheet are produced by a person watching a screen and pressing a key, and the definition behind each key is a decision made by a provider, not a fact about football.
What follows is the sheet read row by row, with the collection method attached to each field and the specific way each one misleads. The aim is not scepticism for its own sake. It is knowing which three rows will support an argument and which twelve will collapse under one.
What the sheet is, and who typed it
- Source. A commercial data provider or the competition’s own operation, working from a broadcast feed or a stadium camera position.
- Method. A mix of automated tracking and human event coding, in proportions the sheet does not disclose.
- Revision. Sheets are corrected after the match. The version circulating on the night is rarely the final one.
- Definitions. Every field is governed by a written definition that is not printed on the sheet. Two providers can produce different numbers and both be correct.
- Trust order. Counts of discrete, unambiguous events first. Judgement-based counts last. Physical data separately, and never across providers.
Reading a football match data sheet from the top down
A football match data sheet rewards a fixed reading order, because several rows only mean something in the presence of another row. Reading it top to bottom in the order printed is how most misreadings begin.
- Header first. Competition, date, venue, and if it is shown, the provider and the revision. A sheet without a provider is a sheet you cannot compare with anything.
- Score and shot line together. Goals, shots, shots on target and blocked shots as one block. Any of them alone is a story; together they are evidence.
- Possession beside pass volume. Never read the possession percentage on its own. Pair it with total passes and the picture changes for about a third of matches.
- Split the pass line into volume and ratio. Accuracy without volume tells you a team passed short and safe, not that it passed well.
- Read the final-third rows. Entries, touches in the box, crosses completed. This is where possession either produced something or did not.
- Read the defensive block as a set. Clearances, interceptions, blocks and aerial duels together describe where a team defended, not how well.
- Discipline line last among the event fields. Fouls and cards are the most context-dependent rows on the sheet.
- Physical data in its own frame. Distance, high-speed running and sprint counts come from a different system and belong to a different conversation.
- Write down what is missing. The absent fields shape the argument as much as the present ones.
Steps two to five take four minutes and produce most of the usable information. The rest of the sheet is context for those rows rather than a second set of conclusions.
Possession is three different numbers wearing one label
There are three common ways to compute possession and they disagree routinely. The first is a clock: an operator holds a key while a team is judged to be in control, and the output is a share of controlled time. The second is a pass share: a team’s passes as a proportion of all passes in the match. The third measures only time while the ball is in play, discarding stoppages.
Those methods produce different answers for the same ninety minutes, and the gap is widest exactly where the number gets quoted most. A team that plays long and direct will look far worse on pass share than on the clock. A team that is repeatedly fouled will look better on the in-play measure than on either alternative.
The practical rule is to treat possession as a description of style rather than of control, and to pair it with two other rows before drawing anything from it: passes attempted, and entries into the final third. A high possession figure with a low final-third entry count is not dominance. It is a team passing in front of a defensive block.
The pass line hides more decisions than any other field
Pass accuracy looks like arithmetic and is actually a definition. Whether a headed clearance that reaches a teammate counts as a completed pass, whether a deflected ball counts as an attempt, whether a goalkeeper’s kick downfield is a pass or a clearance — every one of those is a provider decision, and they move the ratio by several points.
Distance is the second hidden variable. Two teams can post identical accuracy figures while one played almost entirely inside twenty metres and the other attempted forty passes over forty metres. Where a sheet offers pass length bands or forward pass share, those rows are worth more than the headline ratio.
The third problem is that accuracy is scored against the passer regardless of the receiver. A perfectly weighted ball that a striker fails to control is an incomplete pass. Over a season this averages out; over one match it can invert the reading of a midfielder’s afternoon.
Shots on target is a judgment, not a count
Shots are among the more reliable rows because the event is discrete and visible. Shots on target is not, because it requires a coder to decide whether an attempt would have entered the goal without intervention.
Three cases cause most of the disagreement. A shot blocked by a defender is normally excluded even if it was heading in, so a team facing a deep block loses on-target attempts to bodies rather than to accuracy. A shot the goalkeeper gathers comfortably is on target, which puts a soft attempt in the same row as one that required a full-stretch save. And a shot striking the post is counted differently by different providers.
The fix is to read shots on target only alongside blocked shots and total attempts, and to accept that the row measures the goalkeeper’s workload rather than the quality of the finishing.
Duels and tackles are the fields to trust least
Duels are the clearest example of a number that looks objective and is not. A duel is whatever the provider’s definition says it is, and the definitions vary on the essential question of when contact becomes a contest. Aerial duels are counted for both teams, so a match with heavy long-ball play inflates the totals on both sides without telling you anything about who won the ball.
Tackles carry a second problem. Some definitions require the tackler to win possession, others count a tackle that dispossesses without retaining. A defender who wins the ball cleanly and knocks it out for a throw is credited in one system and not in another, and the difference is invisible on the sheet.
There is also a style bias built into the whole family of fields. A team that presses high and wins the ball before a contest develops records fewer tackles than a team that defends deep and tackles inside its own third. Tackle counts describe where the defending happened, not how well it went. This is the field where independent agreement between two coders is weakest, and it is worth reading a general treatment of inter-rater reliability before quoting any of it.
Clearances and interceptions reward the team under pressure
Clearances are counted, not judged, which makes them reliable as a count and useless as a compliment. A high clearance total means a team spent the match defending its own box. It appears most often next to a defeat.
Interceptions are more interesting and more slippery. The boundary between an interception and a block is a coder’s decision, and the boundary between an interception and a loose ball recovered is another. Providers publish separate rows for recoveries in some sheets and merge them in others.
Read the four defensive rows as a shape rather than a score. High clearances with low interceptions describes a deep block absorbing crosses. Low clearances with high interceptions and recoveries describes a team defending further forward and stepping into passing lanes. Neither is better; they are different problems the coach chose.
Physical data comes from a different system with different thresholds

Distance covered, high-speed running and sprint counts do not come from the event coder. They come from optical tracking installed at the venue, or from the club’s own wearable units, and the two are not interchangeable.
The decisive detail is the threshold. A sprint is defined as movement above a set speed for a set duration, and providers choose different values. Change the threshold slightly and the sprint count for the same player in the same match moves substantially. High-speed running has the same problem with a different cut-off. Comparing these numbers across providers or across seasons after a system change is a category error, not a fine distinction.
Total distance is the row most often quoted and the least useful of the three. It is dominated by walking and jogging, it varies with match tempo more than with effort, and a player on a losing team chasing the ball for an hour will produce a high figure that reflects the situation rather than the contribution. The clubs investing in their own tracking infrastructure generally do so to escape this row rather than to publish it.
Automated, semi-automated, hand-coded
| Field | How it is produced | Consistency | Main failure mode |
|---|---|---|---|
| Goals, cards, substitutions | Official record | Very high | Timing of the minute recorded |
| Shots, corners, offsides | Hand-coded, unambiguous events | High | Deflections and rebounds double counted |
| Passes and accuracy | Hand-coded with a definition | Moderate to high | Clearances and headers classified as passes |
| Shots on target | Hand-coded judgement | Moderate | Blocked attempts and posts treated differently |
| Duels, tackles, interceptions | Hand-coded judgement | Low | Definition drift and style bias |
| Distance, sprints | Optical or wearable tracking | High within a system | Thresholds differ between systems |
| Expected goals | Model output from coded shots | Depends on the inputs | Inherits every coding error above it |
The column that matters is the third one. An argument built on rows from the top half of that table can survive scrutiny. An argument built on the duel row cannot, however confident the person making it sounds. This grading of sources is the least glamorous part of sports analytics and the part that decides whether the rest of it is worth anything.
Two providers, one match, two sheets
Where a match is covered by more than one operation, the sheets will differ. Possession by a few points, passes by a few dozen, duels by a great deal. Neither sheet is wrong; they are answering slightly different questions with slightly different rules.
This has a hard practical consequence. Any comparison — between two teams, two players, two seasons — has to be made inside a single provider’s data. Mixing sources produces differences that look like findings and are artefacts of definition. The same applies to league averages: an average computed from one provider cannot be used to place a player recorded by another.
The growth of automated collection has narrowed some of these gaps and widened others, and the wider expansion of commercial sports data operations has made it more common, not less, for two versions of the same match to be in circulation.
What to do when a field looks wrong

Start with the definition rather than the number. Most surprising values are correct outputs of a rule the reader did not know about, and the provider’s glossary resolves them in a minute.
Then check the denominator. A ratio expressed without its base is the single most common source of confusion on a sheet, and a strong percentage built on eleven attempts is not a finding.
Then check whether the sheet has been revised. Corrections after the match are routine, particularly for assists, shot classification and the attribution of a deflected goal, and a number that circulated at midnight may not exist by morning.
Finally, watch the clips. A field that still looks wrong after those three checks usually resolves in ninety seconds of video, and the discipline of going back to the footage is what separates analysis from sheet-reading. Domestic clubs building analysis capacity — including the programmes behind national data initiatives in sport — spend most of their early effort on exactly this loop rather than on new metrics.
Frequently Asked Questions
Which fields on the sheet are genuinely reliable?
Goals, cards, substitutions, corners and offsides, because they are drawn from the official record or from unambiguous events. Shots and passes are close behind. Everything involving a contest between two players is the weakest part of the document.
Why do possession figures differ between broadcasters?
Because they use different methods — a control clock, a share of passes, or time with the ball in play. Each is internally consistent and they do not agree with each other, particularly in matches with a direct style or many stoppages.
Can a data sheet tell me who played well?
It can tell you what a player did, and it cannot tell you what he was asked to do. A holding midfielder instructed to hold position will produce a modest sheet in a match he controlled. The sheet supports a judgement; it does not make one.
Are the physical numbers comparable between leagues?
Only if the same tracking system and the same thresholds are used. Distance and sprint definitions vary between providers, so cross-league comparison of those rows is unsafe without knowing both setups.
How much does the sheet change after revision?
Usually a little, occasionally a lot. Shot classification, assists and the treatment of deflections are the fields most often corrected, and expected goals recomputes when any of them move.
What should a club with a small budget prioritise?
One provider used consistently, a written internal glossary, and the habit of checking video before acting on a number. That combination is worth more than a second data source, and it costs staff time rather than money. Sheets from the same source are far more valuable than more sheets.
A data sheet is a witness statement, not a recording. It is worth reading closely for the same reason a witness statement is: because knowing who wrote it, and under what rules, tells you which parts you can build on.
