Standard II — Integrity of Capital Markets Module 1 · 15-20% Weight Lesson 132

📖 检验统计量、p 值与显著性水平

CFA Level I · L132 · Test Statistic, p-Value & Significance Level

定量方法(Quantitative Methods)— 假设检验 · 第 2 课


一、上节课回顾:从假设到判断

在 L131 中,我们学会了设立假设:

H₀: μ = μ₀(现状假设,含等号)
Hₐ: μ ≠ μ₀ / μ > μ₀ / μ < μ₀(研究主张,不含等号)

但设立假设只是第一步。真正的问题是:

🎯 样本数据要跟 H₀ 偏离多少,我们才敢说"这不是巧合,H₀ 真的有问题"?

这就是本节课要回答的三个核心概念: 1. 检验统计量(Test Statistic)——量化"偏离程度" 2. p 值(p-value)——量化"巧合的概率" 3. 显著性水平 α(Significance Level)——"门槛"标准


二、检验统计量(Test Statistic)

2.1 定义

检验统计量 是一个根据样本数据计算出来的数值,用来衡量样本结果与 H₀ 之间的「差异有多大」。

2.2 直观理解

想象你在射箭:

射击场景 假设检验对应
靶心 = H₀ 声称的值(如 μ = 100) μ₀ = 100
你射出的箭 = 样本均值 x̄ 例如 x̄ = 108
箭离靶心的距离 = 检验统计量 z = (108 - 100) / SE
距离越大 → 越怀疑"瞄准器有问题" 统计量越大 → 越怀疑 H₀ 不成立

2.3 通用公式框架

检验统计量 = (样本统计量 - H₀ 假设值) / 标准误

即: Test Statistic = (Sample Statistic - Hypothesized Value) / Standard Error

2.4 最常见的检验统计量

场景 检验统计量 公式
总体方差 σ² 已知 z 统计量 z = (x̄ - μ₀) / (σ / √n)
总体方差 σ² 未知 t 统计量 t = (x̄ - μ₀) / (s / √n)
检验方差 χ² 统计量 χ² = (n-1)s² / σ₀²
检验两个方差比 F 统计量 F = s₁² / s₂²

📌 CFA 一级重点考察 z 检验和t 检验,χ² 和 F 只需了解基本概念。

2.5 检验统计量的核心直觉

检验统计量 ≈ 信号 / 噪声

信号 = x̄ - μ₀(样本均值离 H₀ 假设有多远)
噪声 = 标准误(抽样波动的正常范围)
比值 含义
统计量 ≈ 0 x̄ 很接近 μ₀ → 没啥理由怀疑 H₀
统计量 > 2(或 < -2) x̄ 偏离了大约 2 个标准误 → 开始可疑了
统计量 > 3 偏离了 3 个标准误以上 → 几乎不可能是巧合

三、显著性水平 α(Significance Level)

3.1 定义

显著性水平 α 是研究者在检验前设定的「最大容忍错误拒绝 H₀ 的概率」。它是判断"证据是否足够强"的门槛。

3.2 常见取值

α 含义 场景
0.05(5%) 容忍 5% 概率冤枉 H₀ 🏆 CFA 默认 / 社会科学研究
0.01(1%) 容忍 1% 概率冤枉 H₀ 医学检验、严格风控
0.10(10%) 容忍 10% 概率冤枉 H₀ 探索性分析、样本量小时

3.3 α 决定「拒绝域」

拒绝域 = 如果检验统计量落入这个区域,就拒绝 H₀。

双尾检验(Hₐ: μ ≠ μ₀):
    ┌──────────┬──────────┬──────────┐
    │  拒绝域   │  不拒绝   │  拒绝域   │
    │  α/2     │   H₀    │  α/2    │
    └──────────┴──────────┴──────────┘
             -z_crit              +z_crit

右尾检验(Hₐ: μ > μ₀):
    ┌────────────────────┬──────────┐
    │      不拒绝 H₀      │  拒绝域   │
    │                    │    α     │
    └────────────────────┴──────────┘
                         +z_crit

3.4 关键临界值(CFA 高频考点)

α 双尾临界值 右尾临界值
0.10 ±1.645 +1.282
0.05 ±1.96 +1.645
0.01 ±2.576 +2.326

🔥 这三个 z 值必须记住,考试不提供!


四、p 值(p-value)

4.1 定义

p 值 是「在 H₀ 为真的前提下,观察到当前样本结果(或更极端结果)的概率」。

4.2 通俗版翻译

p 值 = "如果 H₀ 真的是对的,我们纯靠运气抽到像现在这么离谱的样本的概率有多大?"

p 值 解读
p = 0.80 即使 H₀ 为真,也很容易获得现在的样本 → 完全无法拒绝 H₀
p = 0.15 有一定的随机性可能 → 证据偏弱
p = 0.03 在 H₀ 为真下,只有 3% 概率抽到这么极端的样本 → 证据较强
p = 0.001 几乎不可能在 H₀ 为真时出现 → 强烈拒绝 H₀

4.3 p 值与 α 的比较规则

p < α  →  拒绝 H₀ ✅(统计显著)
p ≥ α  →  不拒绝 H₀(统计不显著)

🎯 一句话:p 值越小,拒绝 H₀ 的证据越强。

4.4 p 值 ≠ H₀ 为真的概率!

这是 CFA 考试最喜欢考的陷阱:

❌ 错误理解:「p = 0.03 意味着 H₀ 有 3% 的概率为真」
✅ 正确理解:「p = 0.03 意味着:如果 H₀ 为真,只有 3% 的概率看到这样的数据」

❌ 错误理解:「p = 0.03 意味着我们有 97% 的把握 H₀ 是假的」
✅ 正确理解:「数据与 H₀ 的矛盾程度相当于 0.03,低于 0.05 所以我们拒绝 H₀」

🔑 核心区别:p 值讨论的是数据在 H₀ 下的条件概率,不是H₀ 在给定数据下的条件概率。这两个在数学上完全不同!


五、三种检验方式的 p 值计算

5.1 双尾检验(Hₐ: μ ≠ μ₀)

p-value = 2 × P(Z > |z_obs|)

即:检验统计量绝对值外侧的两个尾部面积之和

例子: z = 2.1,查表得 P(Z > 2.1) = 0.0179 → p = 2 × 0.0179 = 0.0358

5.2 右尾检验(Hₐ: μ > μ₀)

p-value = P(Z > z_obs)

例子: z = 2.1 → p = P(Z > 2.1) = 0.0179

5.3 左尾检验(Hₐ: μ < μ₀)

p-value = P(Z < z_obs)(z_obs 为负时取左侧面积)

例子: z = -2.1 → p = P(Z < -2.1) = 0.0179

📌 注意:同样 |z| = 2.1,双尾 p 值是单尾的两倍!


六、实战案例:投资组合绩效检验

案例背景

某基金经理声称其投资组合年化超额收益(alpha)大于 0。已知: - 该策略过去 36 个月的数据 - 样本月均 alpha = 0.25% - 样本标准差 s = 0.60% - n = 36 - α = 0.05(5% 显著性水平)

第一步:设立假设

H₀: alpha ≤ 0(没有正的超额收益)
Hₐ: alpha > 0(有正的超额收益)→ 右尾检验

第二步:计算检验统计量(t 统计量,σ 未知)

标准误 SE = s / √n = 0.60% / √36 = 0.60% / 6 = 0.10%

t = (x̄ - μ₀) / SE = (0.25% - 0) / 0.10% = 2.50

第三步:计算 p 值

自由度 df = n - 1 = 35
查 t 分布表:t = 2.50, df = 35
→ 单尾 p ≈ 0.0087

第四步:做出决策

p = 0.0087 < α = 0.05
→ 拒绝 H₀ ✅
→ 结论:有充分证据表明该策略存在正的超额收益

第五步:投资实践解读

统计结论 投资含义
拒绝 H₀(alpha ≤ 0) 样本证据支持 alpha > 0
p = 0.0087 如果 alpha 真的 ≤ 0,只有 0.87% 的概率观察到 t ≥ 2.50
⚠️ 但不等于"alpha 一定 > 0" 仍有 Type I 错误风险

七、p 值的常见误区(CFA 高频考点)

误区 1:p 值越小 = 效应量越大 ❌

p 值受样本量影响。样本量足够大时,即使差异微不足道,p 值也会很小。

n x̄ s 解释
36 0.25% 0.60% t = 2.50, 统计显著
10000 0.01% 0.60% 差异更小但 n 极大 → 也可能显著

💡 0.01% 的 alpha 是否真的有经济意义?统计显著 ≠ 经济显著!

误区 2:p > 0.05 = H₀ 为真 ❌

p > 0.05 只能说「数据没有足够证据推翻 H₀」,不等于「H₀ 是真的」。可能样本量太小(检验力不足)。

误区 3:p < 0.05 = 结果很重要 ❌

统计显著 ≠ 实际重要。月收益高 0.001% 在统计上可能显著(样本够大),但在经济上没有意义。

误区 4:p 值就是「H₀ 为真的概率」❌

这是经典贝叶斯与频率学派的混淆。p 值是 P(Data | H₀),不是 P(H₀ | Data)。


八、α 的选择:严谨 vs 实用的平衡

α 越小 后果
更难拒绝 H₀ 减少 Type I 错误(冤枉好人)
但增加 Type II 错误 更容易「漏过」真实效应(放过坏人)

📌 Type I / Type II 错误将在 L133 详细讨论。

行业惯例

行业 常用 α 原因
学术研究 0.05 传统标准
医学临床试验 0.01 或更低 人命关天,宁可保守
量化交易 0.01 ~ 0.05 取决于策略容量和回撤容忍度
探索性分析 0.10 不想错过可能的信号

九、决策流程图

        提出假设 H₀ / Hₐ
              ↓
        选择 α(如 0.05)
              ↓
        收集样本数据
              ↓
     计算检验统计量(z 或 t)
              ↓
         计算 p 值
              ↓
       ┌──────────────┐
       │  p < α 吗?   │
       └──────┬───────┘
         ↙        ↘
       YES         NO
        ↓           ↓
   拒绝 H₀     不拒绝 H₀
   (显著)     (不显著)
        ↓           ↓
   支持 Hₐ     证据不足

十、课堂练习

📝 Part A:基础概念

Q1. 检验统计量的作用是:

A. 直接判断 H₀ 是否为真 B. 量化样本结果与 H₀ 之间的偏离程度 C. 确定显著性水平 α D. 计算样本均值

Q2. 关于 p 值,以下哪项理解是正确的?

A. p 值是 H₀ 为真的概率 B. p 值是 Hₐ 为真的概率 C. p 值是在 H₀ 为真的前提下,观察到当前(或更极端)样本结果的概率 D. p 值越大,拒绝 H₀ 的证据越强

Q3. 若显著性水平 α = 0.05,以下哪个 p 值可以拒绝 H₀?

A. p = 0.10 B. p = 0.06 C. p = 0.049 D. p = 0.50


📝 Part B:计算题

Q4. 已知 H₀: μ = 100,样本均值 x̄ = 106,总体标准差 σ = 15,n = 25。计算 z 统计量:

A. z = 1.00 B. z = 1.50 C. z = 2.00 D. z = 2.50

Q5. 接上题,该检验的 p 值(双尾)最接近:

A. 0.16 B. 0.05 C. 0.025 D. 0.01

Q6. α = 0.05 时,双尾检验的临界 z 值为:

A. ±1.28 B. ±1.645 C. ±1.96 D. ±2.576


📝 Part C:综合判断

Q7. 某分析师做假设检验得到 p = 0.02,正确的说法是:

A. H₀ 有 2% 的概率为真 B. 如果 H₀ 为真,只有 2% 的概率看到这样的数据 C. Hₐ 有 98% 的概率为真 D. 应该接受 H₀

Q8. 以下哪种说法关于 α 是错误的?

A. α 是检验前设定的容忍错误拒绝 H₀ 的最大概率 B. α 越小,越不容易拒绝 H₀ C. α 的选择应在观察数据之后 D. α = 0.05 意味着最多容忍 5% 的 Type I 错误率

Q9. 在 t 检验中,当样本量增大时:

A. 标准误会增大 B. t 分布趋近于标准正态分布 C. 自由度减小 D. p 值一定减小

Q10. 某交易策略回测得到 p = 0.001,以下解读最合理的是:

A. 该策略 99.9% 一定有效 B. 该策略的收益非常可观 C. 如果该策略真的无效,观察到如此极端结果的概率只有 0.1% D. 应该立即投入全部资金


十一、答案与解析

Part A:基础概念

Q1. 答案:B

检验统计量 = (样本统计量 - H₀ 值) / 标准误,量化的是"偏离程度"。它本身不做判断,需要结合 p 值或临界值。

Q2. 答案:C

p 值 = P(观察到当前或更极端结果 | H₀ 为真)。A/B 都把条件概率搞反了,D 方向反了(p 值越小 → 证据越强)。

Q3. 答案:C

p = 0.049 < α = 0.05 → 拒绝 H₀。p = 0.06 > 0.05 → 不拒绝。


Part B:计算题

Q4. 已知 H₀: μ = 100,x̄ = 106,σ = 15,n = 25。 z = (106 - 100) / (15/√25) = 6 / (15/5) = 6 / 3 = 2.00**

Q5. 答案:B(约 0.05)

z = 2.00,查表 P(Z > 2.00) = 0.0228。双尾 p = 2 × 0.0228 = 0.0456 ≈ 0.05。

Q6. 答案:C

α = 0.05,双尾 → 每侧 α/2 = 0.025 → z = ±1.96(必背值)。


Part C:综合判断

Q7. 答案:B

p 值是 P(Data | H₀),不是 P(H₀ | Data)。p = 0.02 = 如果 H₀ 为真,2% 概率看到这样的数据。

Q8. 答案:C

α 必须在检验前设定,不能在看完数据后再选。这是假设检验的基本原则——避免"套利"显著性水平。

Q9. 答案:B

标准误 = s/√n,n 增大 → 标准误减小,不是增大(A 错)。t 分布随自由度增加趋近于正态分布(B 对)。自由度 df = n-1 增大(C 错)。p 值不一定减小,取决于偏差(D 错)。

Q10. 答案:C

p = 0.001 = 如果策略无效,只有 0.1% 概率看到这样的结果。A 犯了"p 值 = H₀ 为真概率"的错误。B 混淆了统计显著与经济显著。D 太绝对,统计显著不意味着稳赚。


十二、本课要点总结

要点 关键内容
检验统计量 量化样本结果与 H₀ 的偏离程度 = (样本统计量 - H₀值) / 标准误
p 值 P(观察到的数据(或更极端)
显著性水平 α 检验前设定的门槛,常用 0.05 / 0.01 / 0.10
决策规则 p < α → 拒绝 H₀;p ≥ α → 不拒绝 H₀
三个 z 临界值 α=0.10→±1.645; α=0.05→±1.96; α=0.01→±2.576(必背!)
统计显著 ≠ 经济显著 p 值小 ≠ 效应量有实际意义,受样本量影响
p 值误区 p 值不是 H₀ 为真的概率,不是效应量指标,不能后设 α

十三、下节预告:L133

下一课将讨论 Type I 错误与 Type II 错误——"我们可能犯哪两种错?它们之间如何权衡?检验的 Power 又是什么?"


📊 本节 CFA 学习量: ~25 分钟阅读 + 10 道练习题

🏆 恭喜!你已完成 L132,明天见 L133!


CFA Level I · Quantitative Methods · Hypothesis Testing · Lesson 2 Generated for Ivan 哥 on 2026-08-07

Quantitative Methods — Hypothesis Testing · Lesson 2


1. Recap from L131: From Hypotheses to Decisions

In L131, we learned how to set up hypotheses:

H₀: μ = μ₀ (status quo hypothesis, includes equality)
Hₐ: μ ≠ μ₀ / μ > μ₀ / μ < μ₀ (research claim, no equality)

But setting hypotheses is only the first step. The real question is:

🎯 How far must the sample data deviate from H₀ before we can say "this is not coincidence — H₀ really has a problem"?

This lesson answers that question through three core concepts: 1. Test Statistic — quantifies the "degree of deviation" 2. p-value — quantifies the "probability of coincidence" 3. Significance Level α — the "threshold" standard


2. Test Statistic

2.1 Definition

A test statistic is a value calculated from sample data that measures how much the sample result differs from what H₀ predicts.

2.2 Intuitive Understanding

Think of archery:

Archery Scenario Hypothesis Testing
Bullseye = H₀ claim (e.g., μ = 100) μ₀ = 100
Your arrow = sample mean x̄ e.g., x̄ = 108
Distance from bullseye = test statistic z = (108 − 100) / SE
Greater distance → suspect sight is off Larger statistic → doubt H₀ more

2.3 General Formula

Test Statistic = (Sample Statistic − Hypothesized Value) / Standard Error

2.4 Common Test Statistics

Scenario Statistic Formula
σ² known z-statistic z = (x̄ − μ₀) / (σ / √n)
σ² unknown t-statistic t = (x̄ − μ₀) / (s / √n)
Testing variance χ² statistic χ² = (n−1)s² / σ₀²
Testing two variances F-statistic F = s₁² / s₂²

📌 CFA Level I focuses on z-test and t-test. χ² and F require only basic conceptual understanding.

2.5 Core Intuition

Test Statistic ≈ Signal / Noise

Signal = x̄ − μ₀ (how far the sample mean is from H₀)
Noise = Standard Error (normal range of sampling variation)
Ratio Meaning
Statistic ≈ 0 x̄ very close to μ₀ → no reason to doubt H₀
Statistic > 2 (or < −2) x̄ deviates by ~2 SE → starts to look suspicious
Statistic > 3 Deviates by more than 3 SE → almost certainly not coincidence

3. Significance Level α

3.1 Definition

The significance level α is the maximum probability of erroneously rejecting H₀ that the researcher is willing to tolerate, set before the test. It is the threshold for "strong enough evidence."

3.2 Common Values

α Meaning Use Case
0.05 (5%) Tolerate 5% chance of wrongly rejecting H₀ 🏆 CFA default / social sciences
0.01 (1%) Tolerate 1% chance Medical trials, strict risk control
0.10 (10%) Tolerate 10% chance Exploratory analysis, small samples

3.3 α Defines the Rejection Region

Two-tailed test (Hₐ: μ ≠ μ₀):
    ┌──────────┬──────────┬──────────┐
    │ Reject   │  Do not  │  Reject  │
    │  α/2     │  reject  │   α/2    │
    └──────────┴──────────┴──────────┘
          −z_crit              +z_crit

Right-tailed test (Hₐ: μ > μ₀):
    ┌────────────────────┬──────────┐
    │   Do not reject H₀ │  Reject  │
    │                    │    α     │
    └────────────────────┴──────────┘
                         +z_crit

3.4 Critical Values (High-Frequency CFA Topic)

α Two-tailed CV Right-tailed CV
0.10 ±1.645 +1.282
0.05 ±1.96 +1.645
0.01 ±2.576 +2.326

🔥 Memorize these three z-values — they are NOT provided on the exam!


4. p-value

4.1 Definition

The p-value is the probability of observing the current sample result (or a more extreme one), given that H₀ is true.

4.2 Plain-English Translation

p-value = "If H₀ is actually correct, what is the probability that we get a sample this extreme purely by chance?"

p-value Interpretation
p = 0.80 Very easy to get this sample if H₀ true → no reason to reject H₀
p = 0.15 Some randomness possible → weak evidence
p = 0.03 Only 3% chance under H₀ → fairly strong evidence
p = 0.001 Almost impossible under H₀ → strongly reject H₀

4.3 Decision Rule

p < α  →  Reject H₀ ✅ (statistically significant)
p ≥ α  →  Do not reject H₀ (not statistically significant)

🎯 In one line: The smaller the p-value, the stronger the evidence against H₀.

4.4 p-value ≠ Probability that H₀ is true!

This is the CFA exam's favorite trap:

❌ Wrong: "p = 0.03 means there is a 3% probability H₀ is true"
✅ Right: "p = 0.03 means: if H₀ is true, there is only a 3% chance of seeing data like this"

❌ Wrong: "p = 0.03 means we are 97% confident H₀ is false"
✅ Right: "The data contradicts H₀ to a degree of 0.03, below 0.05, so we reject H₀"

🔑 Core difference: p-value = P(Data | H₀), NOT P(H₀ | Data). These are mathematically different!


5. Computing p-values for Three Test Types

5.1 Two-Tailed Test (Hₐ: μ ≠ μ₀)

p-value = 2 × P(Z > |z_obs|)

Sum of both tail areas beyond the absolute value of the test statistic

Example: z = 2.1, table gives P(Z > 2.1) = 0.0179 → p = 2 × 0.0179 = 0.0358

5.2 Right-Tailed Test (Hₐ: μ > μ₀)

p-value = P(Z > z_obs)

Example: z = 2.1 → p = P(Z > 2.1) = 0.0179

5.3 Left-Tailed Test (Hₐ: μ < μ₀)

p-value = P(Z < z_obs)  (take left tail when z_obs is negative)

Example: z = −2.1 → p = P(Z < −2.1) = 0.0179

📌 Note: For the same |z| = 2.1, the two-tailed p-value is twice the one-tailed p-value!


6. Real-World Case: Portfolio Performance Test

Background

A fund manager claims their portfolio generates positive alpha (excess return). Given: - 36 months of historical strategy data - Sample mean monthly alpha = 0.25% - Sample standard deviation s = 0.60% - n = 36 - α = 0.05 (5% significance level)

Step 1: Set Hypotheses

H₀: alpha ≤ 0 (no positive excess return)
Hₐ: alpha > 0 (positive excess return exists) → right-tailed test

Step 2: Compute Test Statistic (t-statistic, σ unknown)

SE = s / √n = 0.60% / √36 = 0.60% / 6 = 0.10%

t = (x̄ − μ₀) / SE = (0.25% − 0) / 0.10% = 2.50

Step 3: Compute p-value

df = n − 1 = 35
t = 2.50 with df = 35
→ one-tailed p ≈ 0.0087

Step 4: Make Decision

p = 0.0087 < α = 0.05
→ Reject H₀ ✅
→ Conclusion: Sufficient evidence of positive alpha

Step 5: Investment Interpretation

Statistical Conclusion Investment Meaning
Reject H₀ (alpha ≤ 0) Sample evidence supports alpha > 0
p = 0.0087 If alpha truly ≤ 0, only 0.87% chance of t ≥ 2.50
⚠️ Does NOT prove alpha > 0 Type I error risk remains

7. Common p-value Misconceptions (High-Frequency CFA Topic)

Misconception 1: Small p-value = Large effect size ❌

p-value is affected by sample size. With a large enough sample, even trivial differences can produce small p-values.

n x̄ s Explanation
36 0.25% 0.60% t = 2.50, significant
10,000 0.01% 0.60% Even smaller difference; large n may still make it significant

💡 Is 0.01% alpha economically meaningful? Statistical significance ≠ Economic significance!

Misconception 2: p > 0.05 means H₀ is true ❌

p > 0.05 only means "insufficient evidence to overturn H₀" — not that H₀ is true. The sample might just be too small (insufficient test power).

Misconception 3: p < 0.05 means the result is important ❌

Statistical significance ≠ practical importance. A 0.001% monthly return could be statistically significant (with huge n) but economically meaningless.

Misconception 4: p-value = Probability that H₀ is true ❌

This is the classic confusion between frequentist and Bayesian thinking. p-value = P(Data | H₀), not P(H₀ | Data).


8. Choosing α: Rigor vs. Practicality

Smaller α Consequence
Harder to reject H₀ Reduces Type I error (false conviction)
But increases Type II error More likely to miss real effects (let the guilty go free)

📌 Type I / Type II errors will be discussed in detail in L133.

Industry Conventions

Industry Common α Reason
Academic research 0.05 Traditional standard
Clinical trials 0.01 or lower Lives at stake; err on the conservative side
Quantitative trading 0.01 ~ 0.05 Depends on strategy capacity and drawdown tolerance
Exploratory analysis 0.10 Do not want to miss potential signals

9. Decision Flowchart

      Formulate H₀ / Hₐ
              ↓
        Choose α (e.g., 0.05)
              ↓
        Collect sample data
              ↓
    Compute test statistic (z or t)
              ↓
         Compute p-value
              ↓
       ┌──────────────┐
       │   p < α ?    │
       └──────┬───────┘
         ↙        ↘
       YES         NO
        ↓           ↓
    Reject H₀   Do not reject H₀
   (significant) (not significant)
        ↓           ↓
    Support Hₐ   Insufficient evidence

10. Practice Questions

📝 Part A: Core Concepts

Q1. The role of a test statistic is to:

A. Directly determine whether H₀ is true B. Quantify the degree of deviation between the sample result and H₀ C. Determine the significance level α D. Calculate the sample mean

Q2. Which of the following is correct about the p-value?

A. The p-value is the probability that H₀ is true B. The p-value is the probability that Hₐ is true C. The p-value is the probability of observing the current (or more extreme) sample result, given that H₀ is true D. The larger the p-value, the stronger the evidence to reject H₀

Q3. With α = 0.05, which p-value leads to rejection of H₀?

A. p = 0.10 B. p = 0.06 C. p = 0.049 D. p = 0.50


📝 Part B: Calculations

Q4. Given H₀: μ = 100, x̄ = 106, σ = 15, n = 25. Compute the z-statistic:

A. z = 1.00 B. z = 1.50 C. z = 2.00 D. z = 2.50

Q5. Continuing from Q4, the two-tailed p-value is closest to:

A. 0.16 B. 0.05 C. 0.025 D. 0.01

Q6. At α = 0.05, the two-tailed critical z-value is:

A. ±1.28 B. ±1.645 C. ±1.96 D. ±2.576


📝 Part C: Integrated Judgment

Q7. An analyst obtains p = 0.02. Which statement is correct?

A. There is a 2% probability that H₀ is true B. If H₀ is true, there is only a 2% chance of seeing such data C. There is a 98% probability that Hₐ is true D. We should accept H₀

Q8. Which statement about α is WRONG?

A. α is the maximum tolerable probability of wrongly rejecting H₀, set before the test B. The smaller α is, the harder it is to reject H₀ C. α should be chosen after observing the data D. α = 0.05 means accepting at most a 5% Type I error rate

Q9. In a t-test, as sample size increases:

A. The standard error increases B. The t-distribution approaches the standard normal distribution C. The degrees of freedom decrease D. The p-value always decreases

Q10. A backtest yields p = 0.001 for a trading strategy. The most reasonable interpretation is:

A. The strategy is 99.9% certain to be effective B. The strategy's returns are very substantial C. If the strategy is truly ineffective, the probability of observing such an extreme result is only 0.1% D. You should immediately invest all your capital


11. Answers and Explanations

Part A: Core Concepts

Q1. Answer: B

Test statistic = (sample statistic − H₀ value) / SE. It quantifies deviation, not decision. Decisions require p-value or critical value comparison.

Q2. Answer: C

p-value = P(observed or more extreme result | H₀ true). A and B reverse the conditional probability. D reverses the direction (smaller p-value → stronger evidence).

Q3. Answer: C

p = 0.049 < α = 0.05 → reject H₀. p = 0.06 > 0.05 → do not reject.


Part B: Calculations

Q4. Answer: C

z = (106 − 100) / (15/√25) = 6 / (15/5) = 6 / 3 = 2.00

Q5. Answer: B (~0.05)

z = 2.00, P(Z > 2.00) = 0.0228. Two-tailed p = 2 × 0.0228 = 0.0456 ≈ 0.05.

Q6. Answer: C

α = 0.05, two-tailed → each tail α/2 = 0.025 → critical z = ±1.96 (must memorize!).


Part C: Integrated Judgment

Q7. Answer: B

p-value = P(Data | H₀), NOT P(H₀ | Data). p = 0.02 = if H₀ true, 2% chance of seeing such data.

Q8. Answer: C

α must be set before the test, not after observing the data. This is a fundamental principle of hypothesis testing — prevents "significance level arbitrage."

Q9. Answer: B

SE = s/√n, as n ↑ → SE ↓, not ↑ (A wrong). t-distribution → normal as df increases (B right). df = n−1 increases (C wrong). p-value may not decrease; depends on deviation (D wrong).

Q10. Answer: C

p = 0.001 = if the strategy is truly ineffective, only 0.1% chance of such an extreme result. A confuses p-value with P(H₀). B confuses statistical and economic significance. D is too absolute; significance does not guarantee profits.


12. Key Takeaways

Key Point Summary
Test Statistic Quantifies deviation from H₀ = (Sample Stat − H₀ value) / SE
p-value P(observed data or more extreme
Significance Level α Threshold set before the test; common values: 0.05, 0.01, 0.10
Decision Rule p < α → Reject H₀; p ≥ α → Do not reject H₀
3 Critical z-values α=0.10→±1.645; α=0.05→±1.96; α=0.01→±2.576 (MUST memorize!)
Stat Sig ≠ Econ Sig Small p-value ≠ meaningful effect size; affected by sample size
p-value Pitfalls p-value ≠ P(H₀ true), ≠ effect size indicator, α cannot be chosen post-hoc

13. Next Lesson: L133

Next up: Type I Error, Type II Error & Power of a Test — "What two kinds of mistakes can we make? How do we balance them? And what is test power?"


📊 Study Load: ~25 min reading + 10 practice questions

🏆 Congratulations! You've completed L132. See you tomorrow for L133!


CFA Level I · Quantitative Methods · Hypothesis Testing · Lesson 2 Generated for Ivan on 2026-08-07

🔜 下一课 · L133

CFA 一级 · L133 · Type I 与 Type II 错误 — 一、上节课回顾:你已经有了判断工具 · 二、四种可能的结果:2×2 决策矩阵 · 三、Type I 错误(α 错误)——「宁可错杀」