本文目录导读:

这是一个很典型的Python数据分析案例复盘场景,下面我按“假设我们有一份球队赛季数据”的思路,把这次分析从背景、方法、代码到结论完整串一遍,重点回答:这次伤病潮是否拖累了球队?
案例背景
假设我们分析的是某支篮球队/足球队的赛季表现,赛季中期球队遭遇了一波伤病潮,多名主力球员连续缺阵,外界普遍认为伤病是球队战绩下滑的主要原因。
我们需要用数据回答三个问题:
- 伤病潮期间,球队战绩是否显著下滑?
- 下滑是因为对手更强,还是自身表现变差?
- 伤病球员的缺阵与球队表现之间是否存在量化关系?
数据准备
假设我们有以下数据表:
表1:比赛记录 games_df
| 字段 | 说明 |
|---|---|
| game_id | 比赛ID |
| date | 日期 |
| opponent | 对手 |
| home_away | 主/客 |
| points_for | 我方得分 |
| points_against | 对方得分 |
| win | 是否获胜 |
| injured_count | 本场缺阵的轮换球员数 |
| key_player_missing | 是否有核心球员缺阵 |
表2:伤病记录 injuries_df
| 字段 | 说明 |
|---|---|
| player | 球员 |
| start_date | 伤停开始 |
| end_date | 伤停结束 |
| importance | 球员重要性评分 |
分析思路
import pandas as pd
import numpy as np
import matplotlib.pyplot as plt
import seaborn as sns
from scipy import stats
# 假设已读取数据
# games_df = pd.read_csv('games.csv')
# injuries_df = pd.read_csv('injuries.csv')
划分阶段:伤病潮前 / 中 / 后
# 定义伤病潮时间窗口
injury_start = '2024-01-10'
injury_end = '2024-02-15'
def phase(date):
if date < injury_start:
return '赛前'
elif date <= injury_end:
return '伤病潮'
else:
return '赛后'
games_df['phase'] = games_df['date'].apply(phase)
各阶段战绩对比
summary = games_df.groupby('phase').agg(
场次=('game_id', 'count'),
胜率=('win', 'mean'),
场均得分=('points_for', 'mean'),
场均失分=('points_against', 'mean'),
场均净胜分=('points_for', 'mean')
).assign(场均净胜分=lambda x: x['场均得分'] - x['场均失分'])
print(summary)
假设输出:
| 阶段 | 场次 | 胜率 | 场均得分 | 场均失分 | 场均净胜分 |
|---|---|---|---|---|---|
| 赛前 | 30 | 667 | 3 | 1 | +7.2 |
| 伤病潮 | 15 | 333 | 8 | 6 | -5.8 |
| 赛后 | 20 | 600 | 9 | 4 | +4.5 |
初步看:伤病潮期间胜率从 66.7% 掉到 33.3%,净胜分从 +7.2 变成 -5.8,确实明显下滑。
统计检验:下滑是否显著?
before = games_df[games_df['phase'] == '赛前']['points_for'] - \
games_df[games_df['phase'] == '赛前']['points_against']
during = games_df[games_df['phase'] == '伤病潮']['points_for'] - \
games_df[games_df['phase'] == '伤病潮']['points_against']
t_stat, p_value = stats.ttest_ind(before, during, equal_var=False)
print(f"t={t_stat:.3f}, p={p_value:.4f}")
p < 0.05,说明净胜分的下滑统计显著,不是随机波动。
控制变量:对手强度
伤病潮期间可能刚好遇到强队,需要排除这个干扰。
# 假设有对手赛季胜率作为强度指标
games_df['opp_strength'] = games_df['opponent'].map(opp_win_rate)
# 分组看:对手强度相近时,伤病是否仍影响战绩
games_df['opp_level'] = pd.cut(games_df['opp_strength'],
bins=[0, 0.4, 0.6, 1.0],
labels=['弱', '中', '强'])
pivot = games_df.pivot_table(index='opp_level', columns='phase',
values='win', aggfunc='mean')
print(pivot)
假设输出:
| 对手水平 | 赛前 | 伤病潮 | 赛后 |
|---|---|---|---|
| 弱 | 85 | 60 | 80 |
| 中 | 65 | 30 | 60 |
| 强 | 40 | 10 | 35 |
即使控制对手强度,伤病潮期间胜率依然全面下滑,说明不能全怪赛程。
伤病数量与战绩的关系
# 按缺阵人数分组
games_df['injured_bin'] = pd.cut(games_df['injured_count'],
bins=[-1, 0, 1, 2, 10],
labels=['0人', '1人', '2人', '3人+'])
record_by_injury = games_df.groupby('injured_bin').agg(
场次=('win', 'count'),
胜率=('win', 'mean'),
净胜分=('points_for', 'mean')
)
record_by_injury['净胜分'] = games_df.groupby('injured_bin').apply(
lambda x: (x['points_for'] - x['points_against']).mean()
)
print(record_by_injury)
假设输出:
| 缺阵人数 | 场次 | 胜率 | 净胜分 |
|---|---|---|---|
| 0人 | 28 | 714 | +8.1 |
| 1人 | 18 | 556 | +2.3 |
| 2人 | 12 | 333 | -4.6 |
| 3人+ | 7 | 143 | -9.8 |
结论非常清晰:缺阵人数越多,胜率和净胜分越差,呈明显负相关。
corr = games_df['injured_count'].corr(games_df['points_for'] - games_df['points_against'])
print(f"缺阵人数与净胜分相关系数: {corr:.3f}")
# 假设输出: -0.62
可视化
fig, axes = plt.subplots(1, 2, figsize=(14, 5))
# 图1:各阶段净胜分
summary['场均净胜分'].plot(kind='bar', ax=axes[0], color=['green','red','blue'])
axes[0].set_title('各阶段场均净胜分')
axes[0].axhline(0, color='black', linewidth=0.8)
# 图2:缺阵人数 vs 净胜分
sns.boxplot(x='injured_bin', y=games_df['points_for'] - games_df['points_against'],
data=games_df, ax=axes[1])
axes[1].set_title('缺阵人数与净胜分关系')
axes[1].axhline(0, color='black', linewidth=0.8)
plt.tight_layout()
plt.show()
✅ 伤病潮确实拖累了球队,证据如下:
- 战绩断层:伤病潮期间胜率从 66.7% → 33.3%,净胜分从 +7.2 → -5.8。
- 统计显著:t 检验 p < 0.05,下滑不是偶然波动。
- 控制对手强度后依然下滑:排除“刚好遇到强队”的解释。
- 剂量效应:缺阵人数越多,表现越差,相关系数约 -0.62。
- 核心球员缺阵影响更大:可进一步用
key_player_missing分组验证。
⚠️ 但要避免过度归因:
- 伤病潮期间可能伴随赛程密集、客场多等因素,需进一步多元回归控制。
- 部分下滑可能来自替补经验不足,而非单纯缺人。
- 样本量只有 15 场,置信区间较宽。
📌 建议
# 多元回归:控制对手强度、主客场、背靠背后,看伤病变量的系数 import statsmodels.api as sm X = games_df[['injured_count', 'opp_strength', 'home_away', 'back_to_back']] X = pd.get_dummies(X, drop_first=True) X = sm.add_constant(X) y = games_df['points_for'] - games_df['points_against'] model = sm.OLS(y, X).fit() print(model.summary())
injured_count 的系数依然显著为负,就可以更有把握地说:伤病潮是拖累球队的独立因素,而非表象。
一句话总结
这次伤病潮不是借口,而是数据上可验证的拖累因素:缺阵人数每增加1人,球队净胜分平均下降约3-4分,胜率下降约15-20个百分点,但要想精确归因,还需要控制赛程、对手强度和主客场等变量。
如果你有真实数据,我可以帮你把上面的代码改成可直接跑的版本,并针对你的字段名做适配。