Standard II — Integrity of Capital Markets Module 1 · 15-20% Weight Lesson 127

📖 标准误(Standard Error)

CFA Level 1 · L127 · Standard Error

定量方法(Quantitative Methods)— 抽样与估计 · 第 3 课


一、从一个问题出发

承接 L126:我们知道 CLT 告诉我们,样本均值 $\bar{x}$ 的抽样分布近似正态,方差为 $\sigma^2/n$。

但这里有个矛盾:要知道抽样分布的方差,必须先知道总体的标准差 σ。而现实中,σ 通常未知。

于是问题来了: 当我们只能得到一个样本时,如何衡量样本均值 $\bar{x}$ 作为总体均值 μ 估计量的精确度?

答案就是标准误。


二、标准误:核心定义

2.1 定义

标准误(Standard Error,简称 SE): 一个统计量的抽样分布的标准差。

对于样本均值 $\bar{x}$:

$$\text{SE}(\bar{x}) = \frac{\sigma}{\sqrt{n}}$$

其中: - $\sigma$ = 总体标准差 - $n$ = 样本容量

2.2 标准误 vs 标准差:关键区别

概念 符号 描述什么 受样本量影响?
总体标准差 $\sigma$ 总体中个体之间的离散程度 ❌ 不变
标准误 $\sigma/\sqrt{n}$ 样本均值作为估计量的精确度 ✅ n 越大越小

🔥 一句话区分: 标准差衡量"数据点有多散";标准误衡量"样本均值估计得有多准"。

2.3 直觉理解

假设总体标准差 σ = 10:

样本量 n 标准误 σ/√n 含义
4 10/2 = 5.0 样本均值误差较大
25 10/5 = 2.0 精度明显提升
100 10/10 = 1.0 精度大幅提升
400 10/20 = 0.5 精度翻倍需要 4 倍样本量

💡 重要规律: 标准误与 √n 成反比。要"精度翻倍"(标准误减半),需要样本量增大到 4 倍。


三、实践中:用样本标准差估计标准误

3.1 现实困境

真实世界中,总体标准差 σ 几乎总是未知的。我们只能用一个样本的标准差 s 来估计它:

$$s = \sqrt{\frac{\sum_{i=1}^{n}(X_i - \bar{x})^2}{n-1}}$$

由此得到标准误的估计值:

$$\widehat{\text{SE}}(\bar{x}) = \frac{s}{\sqrt{n}}$$

⚠️ 注意区别:$\frac{\sigma}{\sqrt{n}}$ 是真实标准误(理论值);$\frac{s}{\sqrt{n}}$ 是估计标准误(实践中使用)。

3.2 实例计算

案例: 你从某基金的历史月度收益率中随机抽取 36 个月,样本均值 $\bar{x}$ = 0.8%,样本标准差 s = 2.4%。

计算样本均值标准误的估计值:

$$\widehat{\text{SE}}(\bar{x}) = \frac{s}{\sqrt{n}} = \frac{2.4\%}{\sqrt{36}} = \frac{2.4\%}{6} = 0.4\%$$

解读: 用这 36 个月数据估计的月均收益率 0.8%,其标准误约 0.4%。这意味着如果我们反复抽 36 个月样本,样本均值在 0.8% ±(约 0.4%)范围波动的概率约 68%(基于 t 分布的近似)。


四、标准误与置信区间

4.1 为什么标准误重要?

标准误是构建置信区间(Confidence Interval)的核心组件。

对于大样本(n ≥ 30),μ 的 95% 置信区间近似为:

$$\bar{x} \pm 1.96 \times \text{SE}(\bar{x})$$

标准误越小 → 置信区间越窄 → 估计越精确。

4.2 接续上例

基金月收益率:$\bar{x}$ = 0.8%,SE ≈ 0.4%

置信水平 z 值 置信区间 宽度
90% 1.645 0.8% ± 0.658% 1.316%
95% 1.96 0.8% ± 0.784% 1.568%
99% 2.576 0.8% ± 1.030% 2.060%

💡 置信水平越高 → 区间越宽 → 需要更大的 SE 倍数。但无论置信水平如何,SE 的大小决定了"基准精度"。


五、决定标准误大小的因素

5.1 两个决定因素

$$\text{SE}(\bar{x}) = \frac{\sigma}{\sqrt{n}}$$

因素 方向 如何降低 SE?
总体变异性 σ σ 越小,SE 越小 无法控制(总体固有属性)
样本容量 n n 越大,SE 越小 增大样本量(唯一可控因素)

5.2 边际收益递减

  • n 从 25 → 100(4 倍):SE 从 σ/5 → σ/10,减半 ✅
  • n 从 100 → 400(4 倍):SE 从 σ/10 → σ/20,再减半 ✅
  • n 从 400 → 625(1.56 倍):SE 从 σ/20 → σ/25,仅减少 20% ⚠️

📉 样本量增加到一定程度后,继续增大带来的精度提升越来越小,成本却线性增长。这就是抽样成本的精算思维。

5.3 实战应用:基金尽调

场景: 你在评估一位基金经理的选股能力。他管理的多空组合月度 alpha(α)均值为 0.35%(年化约 4.2%),但标准差 s = 2.1%,仅有 24 个月的业绩数据。

SE = 2.1% / √24 ≈ 0.429%

95% 置信区间:0.35% ± 1.96 × 0.429% = [-0.49%, 1.19%]

🔴 结论: 区间跨过零——在统计上,不能排除 alpha 为负的可能性。24 个月数据不够!这解释了为什么机构尽调通常要求 3-5 年以上业绩记录。


六、常见混淆与考试陷阱

6.1 "标准差" vs "标准误" 的表述

在 CFA 考试中:

表述 指什么
"The standard deviation of the sample mean" 标准误 SE(= σ/√n)
"The standard deviation of the population" 总体标准差 σ
"The standard deviation of the sample" 样本标准差 s

6.2 错误直觉

❌ "标准误就是样本的标准差除以 √n,所以样本标准差越大标准误越大"

✅ 样本标准差 s 是 σ 的估计,σ 越大 SE 确实越大。但关键是:标准误衡量的不是你数据本身有多散,而是你的估计有多不稳定。

6.3 CLT 修正

CLT 告诉我们:当总体方差已知时,$\bar{x} \sim N(\mu, \sigma^2/n)$。

当总体方差未知(用 s² 估计)时,标准化后的统计量服从t 分布而非正态分布:

$$t = \frac{\bar{x} - \mu}{s/\sqrt{n}} \sim t_{n-1}$$

这就是后续课程会深入讨论的 t 检验基础。


七、本节要点总结

# 要点
1 标准误 = 统计量抽样分布的标准差。样本均值的 SE = σ/√n
2 标准误 ≠ 标准差:标准差描述数据离散,标准误描述估计精度
3 实践中 σ 未知,用 s/√n 估计标准误
4 标准误是构建置信区间的核心组件:CI = $\bar{x}$ ± 临界值 × SE
5 标准误与 √n 成反比,精度翻倍需样本量 ×4
6 增大样本量是降低标准误的唯一可控手段

八、测试题

Q1: 分析师从总体中随机抽取容量为 100 的样本。样本均值 = 50,样本标准差 = 10。样本均值的标准误估计值是多少?

A. 0.1 B. 1.0 C. 10.0


Q2: 某总体标准差为 8。要使得样本均值的标准误不超过 2,最小样本容量应为多少?

A. 4 B. 16 C. 64


Q3: 以下哪种变化最有效地降低样本均值的标准误?

A. 将样本量从 50 增加到 100 B. 将样本量从 100 增加到 200 C. 将样本量从 25 增加到 100


Q4: 分析师报告:"基于 49 个观测值,该策略季度平均超额收益为 1.5%,标准差为 3.5%。" 该均值估计的标准误最接近:

A. 0.071% B. 0.5% C. 7.0%


Q5(判断): 总体服从正态分布时,样本均值的标准误 = σ/√n 这个公式才成立。


答案与解析

Q1:B SE = s/√n = 10/√100 = 10/10 = 1.0

Q2:B n ≥ (σ/SE)² = (8/2)² = 16

Q3:C SE 与 √n 成反比,比较 n 增加倍数: - A:100/50 = 2 → SE × 1/√2 ≈ ×0.707 - B:200/100 = 2 → SE × 1/√2 ≈ ×0.707 - C:100/25 = 4 → SE × 1/2 = ×0.5 ✅

Q4:B SE = s/√n = 3.5%/√49 = 3.5%/7 = 0.5%

Q5:错误。 CLT 保证:无论总体分布如何,只要 n 足够大,样本均值的抽样分布都近似正态,标准误公式 σ/√n 始终成立。关键在于"样本容量足够大"而非"总体是正态的"。


🎯 下一课预告:L128 · 置信区间构建 — 从标准误到区间估计

Quantitative Methods — Sampling and Estimation · Lesson 3


1. Starting with a Question

Continuing from L126: We know that the CLT tells us the sampling distribution of the sample mean $\bar{x}$ is approximately normal, with variance $\sigma^2/n$.

But there is a contradiction: To know the variance of the sampling distribution, we must first know the population standard deviation σ. In reality, σ is usually unknown.

So the question is: When we only have a single sample, how do we measure the precision of the sample mean $\bar{x}$ as an estimator of the population mean μ?

The answer is the Standard Error.


2. Standard Error: Core Definition

2.1 Definition

Standard Error (SE): The standard deviation of the sampling distribution of a statistic.

For the sample mean $\bar{x}$:

$$\text{SE}(\bar{x}) = \frac{\sigma}{\sqrt{n}}$$

Where: - $\sigma$ = population standard deviation - $n$ = sample size

2.2 Standard Error vs. Standard Deviation: Key Differences

Concept Symbol What it describes Affected by sample size?
Population Standard Deviation $\sigma$ Dispersion among individual observations in the population ❌ No
Standard Error $\sigma/\sqrt{n}$ Precision of the sample mean as an estimator ✅ Decreases as n increases

🔥 In one sentence: Standard deviation measures "how spread out the data points are"; Standard Error measures "how precise the sample mean estimate is."

2.3 Intuitive Understanding

Suppose the population standard deviation σ = 10:

Sample Size n Standard Error σ/√n Meaning
4 10/2 = 5.0 Large estimation error
25 10/5 = 2.0 Noticeable precision improvement
100 10/10 = 1.0 Substantial precision improvement
400 10/20 = 0.5 Quadruple sample → halve the SE

💡 Key insight: Standard Error is inversely proportional to √n. To "double precision" (halve the SE), you need to quadruple the sample size.


3. In Practice: Estimating Standard Error Using Sample Standard Deviation

3.1 The Real-World Dilemma

In the real world, the population standard deviation σ is almost always unknown. We can only estimate it using the sample standard deviation s:

$$s = \sqrt{\frac{\sum_{i=1}^{n}(X_i - \bar{x})^2}{n-1}}$$

From this, we obtain the estimated standard error:

$$\widehat{\text{SE}}(\bar{x}) = \frac{s}{\sqrt{n}}$$

⚠️ Note the distinction: $\frac{\sigma}{\sqrt{n}}$ is the true standard error (theoretical); $\frac{s}{\sqrt{n}}$ is the estimated standard error (used in practice).

3.2 Worked Example

Case: You randomly sample 36 months of historical monthly returns from a fund. The sample mean $\bar{x}$ = 0.8% and the sample standard deviation s = 2.4%.

Calculate the estimated standard error of the sample mean:

$$\widehat{\text{SE}}(\bar{x}) = \frac{s}{\sqrt{n}} = \frac{2.4\%}{\sqrt{36}} = \frac{2.4\%}{6} = 0.4\%$$

Interpretation: Using these 36 months of data, the estimated average monthly return is 0.8%, with a standard error of approximately 0.4%. This means that if we repeatedly drew 36-month samples, the sample mean would fall within roughly 0.8% ± 0.4% about 68% of the time (based on the t-distribution approximation).


4. Standard Error and Confidence Intervals

4.1 Why Standard Error Matters

Standard Error is a core component for constructing Confidence Intervals (CI).

For large samples (n ≥ 30), the approximate 95% confidence interval for μ is:

$$\bar{x} \pm 1.96 \times \text{SE}(\bar{x})$$

Smaller SE → narrower CI → more precise estimate.

4.2 Continuing the Example

Fund monthly return: $\bar{x}$ = 0.8%, SE ≈ 0.4%

Confidence Level z-value Confidence Interval Width
90% 1.645 0.8% ± 0.658% 1.316%
95% 1.96 0.8% ± 0.784% 1.568%
99% 2.576 0.8% ± 1.030% 2.060%

💡 Higher confidence level → wider interval → need a larger SE multiplier. But regardless of the confidence level, the size of the SE determines the "baseline precision."


5. Factors That Determine the Standard Error

5.1 Two Determinants

$$\text{SE}(\bar{x}) = \frac{\sigma}{\sqrt{n}}$$

Factor Direction How to reduce SE?
Population variability σ Smaller σ → smaller SE Cannot control (inherent population property)
Sample size n Larger n → smaller SE Increase sample size (the only controllable factor)

5.2 Diminishing Marginal Returns

  • n from 25 → 100 (4×): SE from σ/5 → σ/10, halved ✅
  • n from 100 → 400 (4×): SE from σ/10 → σ/20, halved again ✅
  • n from 400 → 625 (1.56×): SE from σ/20 → σ/25, only reduced by 20% ⚠️

📉 Beyond a certain point, further increases in sample size yield progressively smaller precision gains while costs grow linearly. This is the actuarial thinking behind sampling costs.

5.3 Practical Application: Fund Due Diligence

Scenario: You are evaluating a fund manager's stock-picking ability. The manager's long-short portfolio has a mean monthly alpha (α) of 0.35% (annualized ~4.2%), but with a standard deviation s = 2.1%, based on only 24 months of performance data.

SE = 2.1% / √24 ≈ 0.429%

95% CI: 0.35% ± 1.96 × 0.429% = [-0.49%, 1.19%]

🔴 Conclusion: The interval crosses zero — statistically, we cannot rule out the possibility that alpha is negative. 24 months of data is insufficient! This explains why institutional due diligence typically requires 3-5+ years of track record.


6. Common Confusions and Exam Traps

6.1 "Standard Deviation" vs. "Standard Error" Wording

In the CFA exam:

Wording What it refers to
"The standard deviation of the sample mean" Standard Error (= σ/√n)
"The standard deviation of the population" Population standard deviation σ
"The standard deviation of the sample" Sample standard deviation s

6.2 False Intuition

❌ "Standard Error is just the sample standard deviation divided by √n, so a larger sample standard deviation means a larger SE."

✅ The sample standard deviation s is an estimate of σ, and a larger σ does indeed produce a larger SE. But the key point is: Standard Error measures not how spread out your data is, but how unstable your estimate is.

6.3 CLT Refinement

The CLT tells us: when the population variance is known, $\bar{x} \sim N(\mu, \sigma^2/n)$.

When the population variance is unknown (estimated by s²), the standardized statistic follows a t-distribution rather than a normal distribution:

$$t = \frac{\bar{x} - \mu}{s/\sqrt{n}} \sim t_{n-1}$$

This is the foundation of the t-test, which will be explored in depth in subsequent lessons.


7. Key Takeaways

# Point
1 Standard Error = the standard deviation of the sampling distribution of a statistic. For the sample mean, SE = σ/√n
2 SE ≠ Standard Deviation: SD describes data dispersion; SE describes estimation precision
3 In practice, σ is unknown — use s/√n to estimate the standard error
4 SE is the core component for constructing confidence intervals: CI = $\bar{x}$ ± critical value × SE
5 SE is inversely proportional to √n; quadrupling the sample size halves the SE
6 Increasing sample size is the only controllable way to reduce the standard error

8. Practice Questions

Q1: An analyst draws a random sample of size 100 from a population. The sample mean is 50 and the sample standard deviation is 10. What is the estimated standard error of the sample mean?

A. 0.1 B. 1.0 C. 10.0


Q2: A population has a standard deviation of 8. To ensure the standard error of the sample mean does not exceed 2, what is the minimum required sample size?

A. 4 B. 16 C. 64


Q3: Which of the following changes most effectively reduces the standard error of the sample mean?

A. Increasing sample size from 50 to 100 B. Increasing sample size from 100 to 200 C. Increasing sample size from 25 to 100


Q4: An analyst reports: "Based on 49 observations, the strategy's average quarterly excess return is 1.5%, with a standard deviation of 3.5%." The standard error of this mean estimate is closest to:

A. 0.071% B. 0.5% C. 7.0%


Q5 (True/False): The formula for the standard error of the sample mean, SE = σ/√n, is valid only when the population is normally distributed.


Answers & Explanations

Q1: B SE = s/√n = 10/√100 = 10/10 = 1.0

Q2: B n ≥ (σ/SE)² = (8/2)² = 16

Q3: C SE is inversely proportional to √n. Compare the multiples of n increase: - A: 100/50 = 2 → SE × 1/√2 ≈ ×0.707 - B: 200/100 = 2 → SE × 1/√2 ≈ ×0.707 - C: 100/25 = 4 → SE × 1/2 = ×0.5 ✅

Q4: B SE = s/√n = 3.5%/√49 = 3.5%/7 = 0.5%

Q5: False. The CLT guarantees that regardless of the population distribution, as long as n is sufficiently large, the sampling distribution of the sample mean is approximately normal, and the standard error formula σ/√n always holds. The key condition is "sufficiently large sample size," not "normally distributed population."


🎯 Next Lesson Preview: L128 · Confidence Interval Construction — From Standard Error to Interval Estimation

🔜 下一课 · L128

CFA 一级 · L128 · 点估计 vs 区间估计 — 一、引言:回顾与过渡 · 二、点估计(Point Estimate) · 三、区间估计(Confidence Interva