Analysis · Data

xG: what the statistic really tells you,
and where it lies

Expected goals have become the most quoted statistic in modern football — and the most misused. Here is how an xG figure is built, how to read it without fooling yourself, and the four blind spots that almost nobody mentions.

Published 2 September 2026 · 9 min read · Level: intermediate
Tactical view of a football pitch in midnight-blue wireframe, dotted with cyan points of light of different sizes representing shots and their xG value
Every shot in a match becomes a dot, and every dot carries a probability of a goal. That is all xG is.

1. xG in one sentence (and only one)

The xG of a shot is the probability that the shot ends up in the net, estimated from thousands of comparable shots in the past. A shot rated at 0.30 xG means that shots taken in similar conditions were converted roughly 30 times out of 100.

A team's xG for a match is simply the sum of the xG of all its shots. A team on 1.8 xG has created, across all its chances, the equivalent of “1.8 average goals”. It may just as well have scored 4, or 0.

What xG is not

It is not a performance rating, not a prediction of the score, and not a verdict on a player's quality. It is a measure of the quantity and quality of the chances created. Nothing more, nothing less.

Shot map: xG value by position Diagram of a penalty area showing six shooting positions and the approximate xG associated with each one, from 0.02 for a long-range shot to 0.76 for a penalty. 0.76 PENALTY 0.55 0.08 TIGHT ANGLE 0.09 0.06 0.02 LONG-RANGE SHOT
Typical orders of magnitude seen in public xG models. Exact values vary from one provider to another — which is precisely the subject of section 5. The size of each circle is proportional to the square root of the xG.

2. How xG is actually calculated

Contrary to what you often read, there is no mathematical xG formula that you could apply by hand. An xG value is the output of a trained model, usually a logistic regression or a tree-based machine learning model.

The principle is simple. You gather a database of several hundred thousand shots for which three things are known: the characteristics of the shot, and whether or not it ended up in the net. The model learns to connect the one to the other. It is then shown a new shot, and it returns a probability.

The variables that really matter

VariableWeightWhy
Distance to goalVery strongThe dominant factor: the probability collapses very quickly with every metre.
Shot angleVery strongDetermines how much of the goal is visible. A tight angle cancels out the advantage of being close.
Body partStrongA header is converted less often than a shot with the foot from the same distance.
Type of playStrongCounter-attack, set piece, open play: the defence is not set up in the same way.
Preceding assistMediumA cut-back does not create the same situation as a cross in the air.
Defensive pressureVariableNumber and position of defenders. Only the richest models include it.
The detail that changes everything

The last row is the one that sets the models apart. A model that ignores the position of the defenders sees a shot from 8 metres, dead centre, as a big chance — even if three defenders and the goalkeeper are in the path of the ball. This is the main source of discrepancies between providers.

3. Trap no. 1: per-shot xG versus cumulative xG

This is the most widespread misreading, and it literally reverses the conclusion. Two teams can post the same match xG and have produced two opposite things.

Team ATeam B
Match xG1.601.60
Number of shots420
Average xG per shot0.400.08
InterpretationFew chances, but huge onesLots of shots, all of them poor

Team A created four genuine goalscoring situations. Team B shot twenty times from distance without ever getting into the danger zone. The total is identical; the attacking performance is nothing alike.

The habit to adopt: never read a match xG without dividing it by the number of shots. That one division alone puts you ahead of half the analysis you read online.

4. How many matches before you can trust it?

This is the question that almost no guide addresses seriously, and yet it is the one that decides whether your analysis is worth anything.

A team shoots on average between ten and fifteen times per match. A match xG therefore rests on a sample of a dozen or so events, most of which are worth less than 0.10. A single 0.7 xG shot — a penalty, a one-on-one — is on its own the equivalent of seven or eight long-range efforts. In other words, the xG of a single match is dominated by one or two events.

1match: anecdote
5matches: still very noisy
10-20matches: usable trend

The practical consequence is clear. For a single match, xG is useful for describing what happened — “they dominated but didn't make it count”. It is not useful for predicting anything. To estimate a team's true strength, you need a rolling window of ten to twenty recent matches.

The impossible trade-off

Widening the window reduces the noise, but brings in matches that no longer describe the current team: transfers, a change of manager, injuries. There is no perfect number. Choosing a window means striking a balance between too much noise and too much past.

5. Why two websites do not give the same xG

There is no such thing as one xG. There are xG models, and everyone has their own: Opta, StatsBomb, Understat, not to mention the in-house models of clubs and analysis websites. They are trained on different databases, with different variables.

The result: for the same match, differences of 0.2 to 0.4 xG between two providers are routine. In a match where the key question comes down to 0.3 xG, two honest sources can therefore support two opposite conclusions.

What to take away

6. Overperformance: talent or luck?

A team scores 22 goals from 15 xG. Two narratives compete, and both are plausible:

How to decide is a question of time and repetition. Over ten matches, overperformance is almost always noise. Over three consecutive seasons, with the same attacking personnel, it starts to look like a genuine characteristic.

The costly mistake

Systematically betting against a team “that is outperforming its xG” on the assumption that it will regress from the very next match. Regression to the mean is a long-term phenomenon: it tells you nothing about Saturday's match, and it gives you no date.

7. Three solid uses, two uses to avoid

What works

What doesn't work

8. Frequently asked questions

What does xG mean in football?

xG stands for expected goals. It is the probability that a given shot turns into a goal, estimated from thousands of comparable shots in the past. A shot worth 0.30 xG means that shots taken in similar conditions ended up in the net roughly 30 times out of 100. A team's xG for a match is the sum of the xG of all its shots.

How is xG calculated?

A statistical model — logistic regression or machine learning — is trained on a database of several hundred thousand shots whose outcome is known. It learns to link the characteristics of a shot (distance, angle, body part, type of play, defensive pressure) to a probability of scoring. A new shot is then run through this model, which returns a value between 0 and 1. There is no formula you can apply by hand.

How many matches does it take for xG to be reliable?

The xG from a single match tells you almost nothing: a team takes 10 to 14 shots per match, which is a tiny sample dominated by one or two big events. You need at least around ten matches for the trend to emerge from the noise, and common practice is to work with 10 to 20 recent matches.

Why do two websites give different xG figures for the same match?

Because there is no single xG, only xG models. Each provider trains its own on its own data, with its own variables. A model that knows where the defenders and the goalkeeper were at the moment of the shot will not return the same value as a model that only knows the distance and the angle. Differences of 0.2 to 0.4 xG for the same match are common.

Is a team that outperforms its xG strong or lucky?

Both explanations are possible, and it takes perspective to decide. Over a handful of matches, overperformance is almost always variance and tends to correct itself. Over a whole season, and repeatedly, it can reflect genuine finishing quality that the model does not capture, because it does not know who is taking the shot.

xG, applied to today's matches

Every morning, IASHARK calculates probabilities from real data — including xG created and conceded — and explains the reasoning match by match. The match of the day is free, always.

See today's analysis →