Standard II — Integrity of Capital Markets Module 1 · 15-20% Weight Lesson 129

📖 置信区间构建

CFA Level I · L129 · Confidence Interval Construction

定量方法(Quantitative Methods)— 抽样与估计 · 第 5 课


一、引言

前情回顾(L128):我们学过了置信区间的概念、基本公式和解释方式。 置信区间 = 点估计 ± 临界值 × 标准误

本节课,我们把这个框架落地到四种具体场景,逐一掌握每种场景的公式、适用条件和计算细节。


二、场景一:总体均值的置信区间(σ 已知 / 大样本)

2.1 公式

x̄ ± z(α/2) × σ/√n

2.2 适用条件

条件 说明
总体标准差 σ 已知 罕见但考试爱考
或 n ≥ 30(大样本) 中心极限定理兜底,可用 s 替代 σ

2.3 案例

某银行信用卡部调查了 100 名客户的月均消费额,样本均值 x̄ = ¥4,200,历史数据显示总体标准差 σ = ¥800。构造 95% 置信区间。

SE = 800/√100 = 80

95% CI = 4200 ± 1.96 × 80 = 4200 ± 156.8 = [4043.2, 4356.8]

📌 解读: 我们有 95% 的把握认为,全体客户月均消费额在 ¥4,043 到 ¥4,357 之间。

2.4 关键考点:置信水平 ↔ z 值对照

置信水平 α z(α/2)
90% 0.10 1.645
95% 0.05 1.96
99% 0.01 2.576

三、场景二:总体均值的置信区间(σ 未知 + 小样本)

3.1 公式

x̄ ± t(α/2, df) × s/√n, df = n - 1

3.2 t 分布的关键特征

特征 z 分布(标准正态) t 分布
形状 固定,精确的钟形 钟形,但尾部更厚(fat tail)
依赖参数? 无 依赖自由度 df
df → ∞ — 趋近于 z 分布
同置信水平下的临界值 较小 更大(更保守)

💡 直观理解: σ 未知时,我们用 s 估计 σ,相当于又多了一层不确定性。t 分布的厚尾就是对"连 σ 都不知道"的惩罚。

3.3 常用 t 临界值速查表

df t(0.05)(90%) t(0.025)(95%) t(0.005)(99%)
5 2.015 2.571 4.032
10 1.812 2.228 3.169
15 1.753 2.131 2.947
20 1.725 2.086 2.845
30 1.697 2.042 2.750
∞ 1.645 1.96 2.576

🔥 记忆技巧:df 越小,t 和 z 差距越大;df ≥ 30,t ≈ z(近似)

3.4 案例:小样本下的 CI

分析师随机选取了 16 只新能源股票,计算其过去一年收益率的均值为 12%,样本标准差 s = 8%。请构造 95% 置信区间。

df = 16 - 1 = 15, t(0.025, 15) = 2.131

SE = 8%/√16 = 2%

95% CI = 12% ± 2.131 × 2% = 12% ± 4.262% = [7.738%, 16.262%]

⚠️ 如果用 z 值 1.96 近似计算,CI = [8.08%, 15.92%] —— 区间偏窄,不够保守!


四、场景三:总体比例的置信区间

4.1 公式

p̂ ± z(α/2) × √(p̂(1-p̂)/n)

其中 p̂ = 样本比例(成功数/样本量)

4.2 适用条件

条件 说明
n·p̂ ≥ 10 成功的期望数足够
n·(1-p̂) ≥ 10 失败的期望数足够
用 z 临界值 比例问题始终用 z(不用 t)

⚠️ 如果上述条件不满足(即 n·p̂ < 10 或 n·(1-p̂) < 10),需要用精确方法(如二项分布),但一级不考。

4.3 案例:市场调研

某电商平台随机调查了 400 名用户,其中 260 人对新界面表示满意。构造 90% 置信区间。

p̂ = 260/400 = 0.65

SE = √(0.65 × 0.35 / 400) = √(0.00056875) ≈ 0.02385

90% CI = 0.65 ± 1.645 × 0.02385 = 0.65 ± 0.0392 = [0.6108, 0.6892]

📌 解读: 我们有 90% 的把握认为,整体用户满意度在 61.1% 到 68.9% 之间。


五、场景四:两总体均值之差的置信区间

5.1 为什么需要它?

点估计只能告诉你 x̄₁ - x̄₂ = 3.5,但 3.5 是否"明显不同于 0"?区间估计告诉我们答案。

5.2 两种子场景

子场景 A:独立样本,两总体方差已知(或大样本)

(x̄₁ - x̄₂) ± z(α/2) × √(σ₁²/n₁ + σ₂²/n₂)

子场景 B:独立样本,两总体方差未知但假设相等(pooled)

(x̄₁ - x̄₂) ± t(α/2, n₁+n₂-2) × sp × √(1/n₁ + 1/n₂)

其中 sp² = ((n₁-1)s₁² + (n₂-1)s₂²) / (n₁ + n₂ - 2)(pooled variance)

5.3 案例:新药 vs 安慰剂

临床试验:新药组 50 人,平均血压下降 x̄₁ = 12 mmHg,s₁ = 5 mmHg 安慰剂组 50 人,平均下降 x̄₂ = 4 mmHg,s₂ = 4 mmHg 构造 95% 置信区间(假设方差相等)。

sp² = (49×25 + 49×16) / 98 = 2009/98 ≈ 20.5, sp ≈ 4.527

SE = 4.527 × √(1/50 + 1/50) = 4.527 × 0.2 = 0.9054

df = 50 + 50 - 2 = 98, t(0.025, 98) ≈ 1.984

95% CI = (12 - 4) ± 1.984 × 0.9054 = 8 ± 1.796 = [6.20, 9.80]

🟢 解读: 区间全部 > 0,说明在 95% 置信水平下,新药的降压效果确实优于安慰剂,差异约为 6.2 ~ 9.8 mmHg。


六、影响置信区间宽度的因素(全场景总结)

CI 宽度 ∝ 临界值 × 波动性/√n

因素 变化 宽度 直观理解
样本量 n ↑ ↓ 更多信息 → 更精确
置信水平 (1-α) ↑ ↑ 更高把握 → 更保守
总体波动 σ ↑ ↑ 数据本身噪声大
z → t(小样本) — ↑ t 尾部厚,更保守

🔥 考试重点: 给定一组条件变化,判断 CI 变宽还是变窄。


七、常见陷阱与考试套路

7.1 选错临界值

情况 用 z 还是 t?
σ 已知 z
σ 未知,n ≥ 30 z(近似)
σ 未知,n < 30 t(df = n-1)
比例问题 始终用 z
方差/标准差 用 χ²(一级了解即可)

7.2 混淆 SE 的公式

场景 SE 公式
单个均值 σ/√n
单个比例 √(p̂(1-p̂)/n)
两均值之差(独立) √(σ₁²/n₁ + σ₂²/n₂)

7.3 置信区间的判断技巧

如果 95% CI 包含零(或比例的 0.5 等参考值)→ 该参数与参考值之间无统计学显著差异。

如果 95% CI 全部为正 → 参数显著 大于 参考值。

如果 95% CI 全部为负 → 参数显著 小于 参考值。


八、实战综合案例

案例:基金经理的选股能力评估

某基金经理声称其选股能产生超额收益(alpha)。你收集了其 25 个季度的 alpha 数据: - 均值 x̄ = 0.8%/季度 - 标准差 s = 1.5%/季度

问题 1: 构造 95% 置信区间,判断 alpha 是否显著大于零?

问题 2: 如果 n = 100(其他不变),结论会变吗?

解答 1:

df = 24, t(0.025, 24) ≈ 2.064

SE = 1.5%/√25 = 0.3%

95% CI = 0.8% ± 2.064 × 0.3% = 0.8% ± 0.619% = [0.181%, 1.419%]

🟢 CI 全部为正,alpha 显著 > 0。基金经理的选股能力有统计证据。

解答 2(快速版):

SE = 1.5%/√100 = 0.15%

95% CI ≈ 0.8% ± 1.96 × 0.15% = [0.506%, 1.094%]

CI 更窄,结论更稳健。这就是为什么机构投资者重视长业绩记录(大 n = 小 CI = 强结论)。


九、本节要点总结

# 要点
1 均值 CI(σ 已知):x̄ ± z(α/2) · σ/√n
2 均值 CI(σ 未知小样本):x̄ ± t(α/2, n-1) · s/√n
3 比例 CI:p̂ ± z(α/2) · √(p̂(1-p̂)/n),条件:n·p̂, n·(1-p̂) ≥ 10
4 两均值差 CI:用 pooled SE 或两个独立 SE 加总
5 σ 已知 → z;σ 未知大样本 → z;σ 未知小样本 → t
6 比例问题永远用 z
7 CI 不跨零 → 显著;跨零 → 不显著
8 缩小 CI 最有效的方法:增大 n(唯一可控方式)

十、测试题

Q1: 抽样 36 个数据,样本均值 = 50,总体标准差 σ = 12。总体均值的 90% 置信区间为:

A. [46.71, 53.29] B. [46.08, 53.92] C. [47.00, 53.00]


Q2: 分析师用 n = 16 的样本构造 t 区间,样本均值 = 30,s = 8。自由度为:

A. 16 B. 15 C. 取决于是否知道 σ


Q3: 某民调调查 500 人,300 人支持某政策。支持率的 95% CI 为:

A. [0.556, 0.644] B. [0.557, 0.643] C. [0.560, 0.640]


Q4: 在 σ 未知、n = 10 时,用 z 值代替 t 值构造 CI,会导致:

A. CI 比正确的更宽 B. CI 比正确的更窄 C. CI 不变


Q5: A 组(n₁ = 40, x̄₁ = 82)、B 组(n₂ = 40, x̄₂ = 78),两组方差均已知且相等 σ² = 36。μ₁ - μ₂ 的 95% CI 为:

A. [1.37, 6.63] B. [0.56, 7.44] C. [1.04, 6.96]


Q6(判断): 对于比例估计,如果 n·p̂ = 8(< 10),我们应当用 t 分布代替 z 分布来构造 CI。


答案与解析

Q1:A SE = 12/√36 = 2 90% CI = 50 ± 1.645 × 2 = 50 ± 3.29 = [46.71, 53.29]

Q2:B t 分布的自由度 = n - 1 = 16 - 1 = 15,与是否知道 σ 无关。

Q3:B p̂ = 300/500 = 0.6 SE = √(0.6 × 0.4 / 500) = √0.00048 ≈ 0.0219 95% CI = 0.6 ± 1.96 × 0.0219 = 0.6 ± 0.0429 = [0.5571, 0.6429] 最接近的是 [0.557, 0.643]。

Q4:B t 值(t(0.025,9) ≈ 2.262)> z 值(1.96),所以用 z 会使 CI 比正确值更窄。这在考试中是一个常用陷阱:小样本忘了用 t,CI 偏窄,显得更精确,实则是虚假精确。

Q5:A SE = √(36/40 + 36/40) = √1.8 ≈ 1.3416 95% CI = (82 - 78) ± 1.96 × 1.3416 = 4 ± 2.6295 = [1.37, 6.63]

Q6:错误 比例问题始终用 z 分布。当 n·p̂ < 10 或 n·(1-p̂) < 10 时,应使用精确方法(如基于二项分布),而不是切换为 t 分布。t 分布不适用于比例估计。


Q6:错误 比例问题始终用 z 分布。当 n·p̂ < 10 或 n·(1-p̂) < 10 时,应使用精确方法(如基于二项分布),而不是切换为 t 分布。t 分布不适用于比例估计。


计。


Quantitative Methods — Sampling and Estimation · Lesson 5


I. Introduction

Recap (L128): We learned the concept, basic formula, and interpretation of confidence intervals. CI = Point Estimate ± Critical Value × Standard Error

In this lesson, we apply this framework to four specific scenarios, mastering the formulas, conditions, and calculation details for each.


II. Scenario 1: CI for Population Mean (σ Known / Large Sample)

2.1 Formula

x̄ ± z(α/2) × σ/√n

2.2 Conditions

Condition Note
Population SD σ known Rare in practice, common on exams
Or n ≥ 30 (large sample) CLT kicks in, use s as σ proxy

2.3 Example

A bank surveys 100 customers on monthly spending. Sample mean x̄ = ¥4,200, historical σ = ¥800. Construct a 95% CI.

SE = 800/√100 = 80

95% CI = 4200 ± 1.96 × 80 = 4200 ± 156.8 = [4043.2, 4356.8]

📌 Interpretation: We are 95% confident the population mean monthly spending is between ¥4,043 and ¥4,357.

2.4 Key: Confidence Level ↔ z-Value

Confidence Level α z(α/2)
90% 0.10 1.645
95% 0.05 1.96
99% 0.01 2.576

III. Scenario 2: CI for Population Mean (σ Unknown + Small Sample)

3.1 Formula

x̄ ± t(α/2, df) × s/√n, df = n - 1

3.2 Key Features of t-Distribution

Feature z (Standard Normal) t
Shape Fixed, exact bell Bell-shaped, thicker tails
Depends on parameter? No Depends on df
df → ∞ — Approaches z
Critical value for same CL Smaller Larger (more conservative)

💡 Intuition: When σ is unknown, using s to estimate σ adds another layer of uncertainty. The fat tails of the t-distribution "penalize" us for not knowing σ.

3.3 Quick Reference: t Critical Values

df t(0.05) (90%) t(0.025) (95%) t(0.005) (99%)
5 2.015 2.571 4.032
10 1.812 2.228 3.169
15 1.753 2.131 2.947
20 1.725 2.086 2.845
30 1.697 2.042 2.750
∞ 1.645 1.96 2.576

🔥 Memory tip: Smaller df → larger gap between t and z; df ≥ 30, t ≈ z.

3.4 Example: Small-Sample CI

An analyst randomly selects 16 clean energy stocks. Their past-year mean return is 12%, sample SD s = 8%. Construct a 95% CI.

df = 16 - 1 = 15, t(0.025, 15) = 2.131

SE = 8%/√16 = 2%

95% CI = 12% ± 2.131 × 2% = 12% ± 4.262% = [7.738%, 16.262%]

⚠️ Using z = 1.96 would give CI = [8.08%, 15.92%] — narrower, less conservative!


IV. Scenario 3: CI for Population Proportion

4.1 Formula

p̂ ± z(α/2) × √(p̂(1-p̂)/n)

where p̂ = sample proportion (successes / sample size)

4.2 Conditions

Condition Note
n·p̂ ≥ 10 Sufficient expected successes
n·(1-p̂) ≥ 10 Sufficient expected failures
Use z critical values Proportions ALWAYS use z (never t)

⚠️ If conditions are not met (n·p̂ < 10 or n·(1-p̂) < 10), exact methods (binomial) are required — not tested at Level I.

4.3 Example: Market Survey

An e-commerce platform surveys 400 users; 260 are satisfied with a new UI. Construct a 90% CI.

p̂ = 260/400 = 0.65

SE = √(0.65 × 0.35 / 400) = √(0.00056875) ≈ 0.02385

90% CI = 0.65 ± 1.645 × 0.02385 = 0.65 ± 0.0392 = [0.6108, 0.6892]

📌 Interpretation: We are 90% confident the true satisfaction rate is between 61.1% and 68.9%.


V. Scenario 4: CI for Difference Between Two Population Means

5.1 Why Do We Need This?

A point estimate tells you x̄₁ - x̄₂ = 3.5, but is 3.5 "meaningfully different from 0"? The interval estimate answers this.

5.2 Two Sub-Scenarios

Sub-scenario A: Independent samples, variances known (or large samples)

(x̄₁ - x̄₂) ± z(α/2) × √(σ₁²/n₁ + σ₂²/n₂)

Sub-scenario B: Independent samples, variances unknown but assumed equal (pooled)

(x̄₁ - x̄₂) ± t(α/2, n₁+n₂-2) × sp × √(1/n₁ + 1/n₂)

where sp² = ((n₁-1)s₁² + (n₂-1)s₂²) / (n₁ + n₂ - 2) (pooled variance)

5.3 Example: New Drug vs. Placebo

Clinical trial: Drug group n₁ = 50, mean BP reduction x̄₁ = 12 mmHg, s₁ = 5 mmHg Placebo group n₂ = 50, mean reduction x̄₂ = 4 mmHg, s₂ = 4 mmHg Construct 95% CI (assume equal variances).

sp² = (49×25 + 49×16) / 98 = 2009/98 ≈ 20.5, sp ≈ 4.527

SE = 4.527 × √(1/50 + 1/50) = 4.527 × 0.2 = 0.9054

df = 50 + 50 - 2 = 98, t(0.025, 98) ≈ 1.984

95% CI = (12 - 4) ± 1.984 × 0.9054 = 8 ± 1.796 = [6.20, 9.80]

🟢 Interpretation: CI is entirely > 0, meaning at 95% confidence, the drug truly outperforms placebo, with a reduction difference of ~6.2 to 9.8 mmHg.


VI. Factors Affecting CI Width (All Scenarios)

CI Width ∝ Critical Value × Variability/√n

Factor Change Width Intuition
Sample size n ↑ ↓ More information → more precise
Confidence level (1-α) ↑ ↑ Higher certainty → wider range
Population variability σ ↑ ↑ Noisier data
z → t (small n) — ↑ t has fatter tails, more conservative

🔥 Exam focus: Given a change in conditions, determine whether CI widens or narrows.


VII. Common Pitfalls & Exam Tricks

7.1 Choosing the Wrong Critical Value

Situation z or t?
σ known z
σ unknown, n ≥ 30 z (approximate)
σ unknown, n < 30 t (df = n-1)
Proportions Always z
Variance/SD Use χ² (Level I awareness only)

7.2 Confusing SE Formulas

Scenario SE Formula
Single mean σ/√n
Single proportion √(p̂(1-p̂)/n)
Difference of two means (independent) √(σ₁²/n₁ + σ₂²/n₂)

7.3 CI Judgment Shortcut

If 95% CI includes zero → parameter is NOT significantly different from the reference value.

If 95% CI is entirely positive → parameter is significantly greater than the reference.

If 95% CI is entirely negative → parameter is significantly less than the reference.


VIII. Hands-On Case Study

Case: Evaluating a Fund Manager's Stock-Picking Skill

A fund manager claims to generate alpha. You collect 25 quarters of alpha data: - Mean x̄ = 0.8%/quarter - SD s = 1.5%/quarter

Q1: Construct a 95% CI. Is alpha significantly > 0?

Q2: If n = 100 (everything else unchanged), does the conclusion change?

Solution 1:

df = 24, t(0.025, 24) ≈ 2.064

SE = 1.5%/√25 = 0.3%

95% CI = 0.8% ± 2.064 × 0.3% = 0.8% ± 0.619% = [0.181%, 1.419%]

🟢 CI entirely positive → alpha is significantly > 0. Statistical evidence supports the manager's stock-picking skill.

Solution 2 (quick version):

SE = 1.5%/√100 = 0.15%

95% CI ≈ 0.8% ± 1.96 × 0.15% = [0.506%, 1.094%]

CI is narrower, conclusion is more robust. This is why institutional investors value long track records (large n = narrow CI = strong conclusion).


IX. Key Takeaways

# Takeaway
1 CI for mean (σ known): x̄ ± z(α/2) · σ/√n
2 CI for mean (σ unknown, small n): x̄ ± t(α/2, n-1) · s/√n
3 CI for proportion: p̂ ± z(α/2) · √(p̂(1-p̂)/n); conditions: n·p̂, n·(1-p̂) ≥ 10
4 CI for difference of means: use pooled SE or sum of independent SEs
5 σ known → z; σ unknown + large n → z; σ unknown + small n → t
6 Proportions ALWAYS use z
7 CI does not cross zero → significant; crosses zero → not significant
8 Best way to narrow CI: increase n (the only controllable factor)

X. Practice Questions

Q1: Sample of 36 observations, sample mean = 50, population σ = 12. The 90% CI for the population mean is:

A. [46.71, 53.29] B. [46.08, 53.92] C. [47.00, 53.00]


Q2: An analyst uses n = 16 to construct a t-interval, sample mean = 30, s = 8. Degrees of freedom:

A. 16 B. 15 C. Depends on whether σ is known


Q3: A poll surveys 500 people; 300 support a policy. The 95% CI for the support rate is:

A. [0.556, 0.644] B. [0.557, 0.643] C. [0.560, 0.640]


Q4: When σ is unknown and n = 10, using z instead of t to construct the CI will result in:

A. CI wider than the correct one B. CI narrower than the correct one C. CI unchanged


Q5: Group A (n₁ = 40, x̄₁ = 82), Group B (n₂ = 40, x̄₂ = 78), both variances known and equal σ² = 36. The 95% CI for μ₁ - μ₂ is:

A. [1.37, 6.63] B. [0.56, 7.44] C. [1.04, 6.96]


Q6 (True/False): For proportion estimation, if n·p̂ = 8 (< 10), we should use the t-distribution instead of z to construct the CI.


Answers & Explanations

Q1: A SE = 12/√36 = 2 90% CI = 50 ± 1.645 × 2 = 50 ± 3.29 = [46.71, 53.29]

Q2: B t-distribution df = n - 1 = 16 - 1 = 15, regardless of whether σ is known.

Q3: B p̂ = 300/500 = 0.6 SE = √(0.6 × 0.4 / 500) = √0.00048 ≈ 0.0219 95% CI = 0.6 ± 1.96 × 0.0219 = 0.6 ± 0.0429 = [0.5571, 0.6429] Closest match: [0.557, 0.643].

Q4: B t(0.025, 9) ≈ 2.262 > z = 1.96, so using z produces a CI narrower than the correct one. Classic exam trap: forgetting to use t for small samples makes the CI falsely precise.

Q5: A SE = √(36/40 + 36/40) = √1.8 ≈ 1.3416 95% CI = (82 - 78) ± 1.96 × 1.3416 = 4 ± 2.6295 = [1.37, 6.63]

Q6: False Proportions ALWAYS use z. When n·p̂ < 10 or n·(1-p̂) < 10, use exact methods (e.g., binomial-based), NOT the t-distribution. t is never appropriate for proportion estimation.


🔜 下一课 · L130

CFA 一级 · L130 · 抽样与估计综合练习 — 一、知识回顾:抽样与估计五课速览 · 二、核心公式速查卡 · 三、综合练习题