用 Python 分析足球"边锋内切"打法的威胁
边锋内切(Winger Cutting Inside)是当今足坛非常流行的进攻方式——边锋从边路带球向中路切入,寻找射门或传球机会(如罗本、萨拉赫、马赫雷斯),下面我用 Python 从数据获取 → 特征提取 → 威胁建模 → 可视化的完整流程来演示分析。

整体思路
要量化"内切威胁",核心是把它拆成几个可测量的维度:
| 维度 | 含义 | 数据来源 |
|---|---|---|
| 内切频率 | 单位时间内切向中路的次数 | 事件数据(带球/传球) |
| 内切位置 | 从边路向中路切入的起点 | 坐标 (x, y) |
| 内切后产出 | 射门、助攻、关键传球 | 事件结果 |
| 防守反应 | 吸引几人包夹、是否造犯规 | 事件上下文 |
| xT / xG 增益 | 威胁度变化 | 位置价值模型 |
数据准备
常用数据源:StatsBomb Open Data(免费,含 360 数据)、FBref、Wyscout。
import pandas as pd
import numpy as np
import matplotlib.pyplot as plt
from matplotlib.patches import Arc, Rectangle
from mplsoccer import Pitch
# 读取 StatsBomb 开放数据(示例:某场比赛事件数据)
events = pd.read_json("events/3788741.json")
# 关键字段提取
df = events[[
"id", "player", "type", "location", "pass", "carry",
"shot", "minute", "under_pressure"
]].copy()
df["event_type"] = df["type"].apply(lambda x: x["name"] if isinstance(x, dict) else None)
df["x"] = df["location"].apply(lambda l: l[0] if isinstance(l, list) else None)
df["y"] = df["location"].apply(lambda l: l[1] if isinstance(l, list) else None)
df = df.dropna(subset=["x", "y"])
识别"内切"动作
定义:球员从边路(|y - 40| > 15)向中路移动,且 x 方向推进或横向位移超过阈值。
def detect_cut_inside(df, x_threshold=5, y_shift_threshold=8):
"""
检测内切动作:
- 起点在边路 (y < 25 或 y > 43, 场地宽度 68)
- 带球后有明显的向中路位移
"""
carries = df[df["event_type"] == "Carry"].copy()
carries = carries.sort_values(["player", "minute"])
cut_events = []
for player, group in carries.groupby("player"):
rows = group.to_dict("records")
for i in range(len(rows) - 1):
cur, nxt = rows[i], rows[i + 1]
start_y = cur["y"]
end_y = nxt["y"]
dx = nxt["x"] - cur["x"]
dy = end_y - start_y
# 起点在边路
on_wing = start_y < 25 or start_y > 43
# 向中路移动:左路 y 增大,右路 y 减小
moving_inside = (start_y < 25 and dy > y_shift_threshold) or \
(start_y > 43 and dy < -y_shift_threshold)
# 向前推进
forward = dx > x_threshold
if on_wing and moving_inside and forward:
cut_events.append({
"player": player,
"minute": cur["minute"],
"start_x": cur["x"], "start_y": start_y,
"end_x": nxt["x"], "end_y": end_y,
"dx": dx, "dy": dy,
})
return pd.DataFrame(cut_events)
构建威胁度模型(xT)
用Expected Threat(xT)衡量每次内切前后位置的威胁值变化:
# xT 网格(简化版,实际可用 Karun Singh 公开的 12x8 矩阵)
xt_grid = np.array([
[0.006, 0.007, 0.008, 0.009, 0.012, 0.016, 0.022, 0.030, 0.041, 0.055, 0.070, 0.082],
[0.009, 0.010, 0.011, 0.013, 0.017, 0.022, 0.030, 0.040, 0.053, 0.068, 0.083, 0.093],
# ... 省略其余行,实际需完整 12 行 × 8 列
])
def get_xt(x, y):
"""根据坐标返回 xT 值"""
xi = min(int(x / 105 * 12), 11)
yi = min(int(y / 68 * 8), 7)
return xt_grid[xi, yi]
cuts = detect_cut_inside(df)
cuts["xt_gain"] = cuts.apply(
lambda r: get_xt(r["end_x"], r["end_y"]) - get_xt(r["start_x"], r["start_y"]), axis=1
)
关联内切后的结果(射门/助攻/关键传球)
def get_post_action(player, minute, window=2):
"""取内切后 2 分钟内的产出事件"""
subsequent = df[
(df["player"] == player) &
(df["minute"] >= minute) &
(df["minute"] <= minute + window)
]
outcomes = set()
for _, row in subsequent.iterrows():
if row["event_type"] == "Shot":
outcomes.add("shot")
if row["event_type"] == "Pass" and isinstance(row["pass"], dict):
if row["pass"].get("shot_assist"):
outcomes.add("key_pass")
if row["pass"].get("goal_assist"):
outcomes.add("assist")
return outcomes
cuts["outcomes"] = cuts.apply(
lambda r: get_post_action(r["player"], r["minute"]), axis=1
)
cuts["is_shot"] = cuts["outcomes"].apply(lambda s: "shot" in s)
cuts["is_goal"] = cuts["outcomes"].apply(lambda s: "assist" in s)
可视化:绘制内切轨迹
from mplsoccer import Pitch
# 从 StatsBomb 提取球员全名
player_names = {p["id"]: p["name"] for p in events["player"].dropna().unique()}
pitch = Pitch(pitch_type="statsbomb", line_color="black")
fig, ax = pitch.draw(figsize=(12, 8))
# 按球员绘制,颜色区分
for player, grp in cuts.groupby("player"):
pitch.arrows(
grp["start_x"], grp["start_y"],
grp["end_x"], grp["end_y"],
ax=ax, color="red", width=2, headwidth=8, alpha=0.6,
)
ax.set_title("边锋内切轨迹图", fontsize=16)
plt.savefig("cut_inside.png", dpi=150, bbox_inches="tight")
威胁评估汇总
summary = cuts.groupby("player").agg(
cuts_count=("start_x", "count"),
avg_xt_gain=("xt_gain", "mean"),
total_xt=("xt_gain", "sum"),
shots_after=("is_shot", "sum"),
goals_after=("is_goal", "sum"),
).sort_values("total_xt", ascending=False)
# 威胁指数(可自定义权重)
summary["threat_score"] = (
summary["cuts_count"] * 0.3 +
summary["total_xt"] * 0.5 +
summary["shots_after"] * 0.2
).round(3)
print(summary)
输出示例:
cuts_count avg_xt_gain total_xt shots_after threat_score
Salah, Mohamed 12 0.045 0.54 4 3.87
Mané, Sadio 8 0.038 0.30 2 2.55
...
进阶分析方向
- 防守压力:用
under_pressure字段加权——被逼抢下完成内切,威胁更高。 - 方向偏好:统计左路内切 vs 右路内切的有效性(例如罗本左脚右翼内切)。
- 对手反应:结合 360 数据看内切后对方防线是否被拉扯出空当。
- 机器学习模型:
from xgboost import XGBClassifier # 用内切前的位置、速度、防守人数、比赛状态预测是否形成射门 features = ["start_x", "start_y", "dx", "dy", "under_pressure", "defenders_nearby"] model = XGBClassifier() model.fit(cuts[features], cuts["is_shot"])
- 对比分析:把"内切"与"下底传中"、"禁区外射门"等其他打法做 xT 收益对比。
| 步骤 | 关键点 |
|---|---|
| 数据 | StatsBomb / Wyscout 事件数据 |
| 识别 | 边路起点 + 向中路位移 + 向前推进 |
| 量化 | xT 增益、射门/助攻产出、防守压力 |
| 可视化 | mplsoccer 箭头轨迹 |
| 建模 | 逻辑回归 / XGBoost 预测威胁 |
这样一套流程可以客观地回答"某位边锋的内切是否真的具有威胁",并可以在球队战术分析、对手球探报告中直接使用。
需要我针对某位具体球员(如萨拉赫、罗本)写一个完整可运行的 demo 吗?