Standard II — Integrity of Capital Markets Module 1 · 15-20% Weight Lesson 108

📖 离散程度:极差、MAD、方差、标准差

CFA Level 1 · L108 · Measures of Dispersion: Range, MAD, Variance, Standard Deviation

课题:别被平均值骗了——学会看"离散度"才是真本事


一、引言:两碗汤的故事

两家餐馆,顾客评分平均值都是 4.0 分:

  • 甲餐馆: 评分分布 [4, 4, 4, 4, 4]——所有顾客一致好评
  • 乙餐馆: 评分分布 [1, 1, 4, 5, 9]——有人爱得要死,有人恨之入骨

均值一样,但两家餐馆完全不同。

甲餐馆稳定可预期,乙餐馆像坐过山车。

🔥 核心直觉:均值告诉你"中心在哪",离散程度告诉你"数据散得有多开"。两项必须搭配看,缺一不可。


二、什么是离散程度(Dispersion)?

2.1 定义

离散程度(Dispersion / Variability) 衡量数据围绕中心位置的分散或聚集程度。

  • 离散度小 → 数据紧密聚集在均值附近(低波动)
  • 离散度大 → 数据分布广阔(高波动)

2.2 在投资中的意义

离散度 含义 投资启示
低 收益稳定,波动小 适合保守型投资者
高 收益大起大落,波动大 高风险高回报,需较高风险承受力

📌 离散度是 CFA 中风险度量的数学基础。从方差到贝塔系数再到 VaR,都始于今天这四个工具。


三、四大离散度指标全景

指标 英文 公式思路 单位 对异常值敏感性
极差 Range Max − Min 原始单位 🔴 极度敏感
平均绝对偏差 MAD 平均 |Xᵢ − X̄| 原始单位 🟡 较敏感
方差 Variance 平均 (Xᵢ − μ)² 平方单位 🔴 敏感
标准差 Std Dev √方差 原始单位 🟡 较敏感

四、逐项详解

4.1 极差(Range)

公式:Range = X_max − X_min

优点: 最直观——"最大减最小",秒懂。

致命缺陷: 只看两个极端值,完全忽略中间数据。

实战案例:

基金 A 和基金 B 的月回报: - A:[−2%, 1%, 2%, 3%, 5%] → Range = 5% − (−2%) = 7% - B:[−2%, −1%, 2%, 3%, 12%] → Range = 12% − (−2%) = 14%

只看极差,B 风险更大。但 B 只是 12% 那一个月份异常,其他月份和 A 差不多。

📌 投资界几乎不单独使用极差,因为一个异常值就能让极差失真。

4.2 平均绝对偏差(MAD — Mean Absolute Deviation)

公式:

$$\text{MAD} = \frac{\sum_{i=1}^{n} |X_i - \bar{X}|}{n}$$

计算步骤: 1. 计算均值 X̄ 2. 计算每个值与均值的绝对偏差 |Xᵢ − X̄| 3. 求所有绝对偏差的均值

为什么用绝对值?

因为原始偏差 (Xᵢ − X̄) 的和恒为 0(正负抵消),不用绝对值的话 MAD 永远是 0,没有意义。

实战计算:

数据集:[3, 5, 7, 9, 11]

X̄ = (3+5+7+9+11) / 5 = 7

Xᵢ Xᵢ − X̄ |Xᵢ − X̄|
3 −4 4
5 −2 2
7 0 0
9 +2 2
11 +4 4

MAD = (4+2+0+2+4) / 5 = 12/5 = 2.4

解读: 平均而言,每个数据点距离均值 2.4 个单位。

📌 MAD 好理解、单位不变,但数学上不方便(绝对值不可导),所以方差/标准差后来居上。

4.3 方差(Variance)——离散度的"发动机"

4.3.1 总体方差(Population Variance)

公式:

$$\sigma^2 = \frac{\sum_{i=1}^{N} (X_i - \mu)^2}{N}$$

4.3.2 样本方差(Sample Variance)

公式:

$$s^2 = \frac{\sum_{i=1}^{n} (X_i - \bar{X})^2}{n-1}$$

4.3.3 为什么分母是 n−1 而不是 n?

这是 CFA 一级必考题。

直觉解释: 样本均值 X̄ 是从样本本身算出来的,比总体均值 μ "更靠近"样本数据点,导致偏差 (Xᵢ − X̄) 系统性地偏小。

除以 n−1(而不是 n)就像把蛋糕分给 n−1 个人,每份更大一点——补偿了低估的偏差。

技术术语: 这叫贝塞尔校正(Bessel's Correction),使样本方差成为总体方差的无偏估计量。

🔑 记忆口诀:"总 N 样 N 减一" —— 总体用 N,样本用 n−1。

4.3.4 方差实战计算

数据集(总体):[3, 5, 7, 9, 11],μ = 7

Xᵢ Xᵢ − μ (Xᵢ − μ)²
3 −4 16
5 −2 4
7 0 0
9 +2 4
11 +4 16

σ² = (16+4+0+4+16) / 5 = 40/5 = 8.0

4.3.5 方差的核心缺陷

方差的单位是原始单位的平方。

  • 如果数据是"收益率 %" → 方差单位是 "%²"(百分之平方?)
  • 如果数据是"价格元" → 方差单位是 "元²"

这就是为什么我们需要标准差——把它拉回原始单位。

4.4 标准差(Standard Deviation)——离散度的"最终形态"

公式:

总体标准差:$$\sigma = \sqrt{\sigma^2}$$

样本标准差:$$s = \sqrt{s^2}$$

上例: σ = √8 = 2.828

解读: 在正态分布下,约 68% 的数据落在均值 ± 1 个标准差的范围内。


五、四大指标对比总结

维度 极差 MAD 方差 标准差
使用所有数据? ❌ 仅两个点 ✅ ✅ ✅
单位与原始一致? ✅ ✅ ❌ 平方单位 ✅
数学性质优良? ❌ ⚠️ 绝对值不可导 ✅ 可导 ✅ 可导
对异常值稳健? ❌ ⚠️ ❌ ⚠️
CFA 考试重要性 ⭐ ⭐⭐ ⭐⭐⭐⭐⭐ ⭐⭐⭐⭐⭐

六、切比雪夫不等式(Chebyshev's Inequality)

6.1 定义

对于任何分布的数据,在均值 ± k 个标准差范围内的数据比例至少为:

$$1 - \frac{1}{k^2}, \quad k > 1$$

6.2 应用

k 至少包含比例 计算
1.25 36% 1 − 1/1.5625
1.5 55.6% 1 − 1/2.25
2 75% 1 − 1/4
2.5 84% 1 − 1/6.25
3 88.9% 1 − 1/9

📌 k=2(至少 75%)和 k=3(至少 88.9%)是 CFA 一级高频考点。

6.3 对比:切比雪夫 vs 正态分布

区间 切比雪夫(任意分布) 正态分布
μ ± 1σ 不适用(k≤1) 约 68%
μ ± 2σ 至少 75% 约 95%
μ ± 3σ 至少 89% 约 99.7%

🔥 切比雪夫是"保底"估计——对于任何分布都是安全的。正态分布则给出精确比例。


七、变异系数(Coefficient of Variation, CV)

7.1 为什么需要 CV?

问题:A 基金平均收益 12%,标准差 8%;B 基金平均收益 5%,标准差 4%。谁的风险更大(单位收益承担了多少波动)?

单独的 σ 无法比较——两只基金均值不同,标准差尺度天然不同。

7.2 公式

$$\text{CV} = \frac{s}{\bar{X}} \quad \text{或} \quad \frac{\sigma}{\mu}$$

7.3 实战计算

  • A 基金:CV_A = 8% / 12% = 0.667
  • B 基金:CV_B = 4% / 5% = 0.800

B 基金的 CV 更大 → 每单位收益承受的波动更大。

7.4 CV 的局限

⚠️ 当均值接近 0 或为负数时,CV 失效(除以接近 0 的值会产生荒谬结果)。


八、实战应用:投资组合风险识别

场景

你正在分析三只股票的年化日波动率数据(已计算标准差):

股票 均值回报 标准差 CV
X 15% 12% 0.80
Y 8% 8% 1.00
Z 20% 18% 0.90

分析: - 单独看标准差:Z 风险最大(σ=18%),Y 风险最小(σ=8%) - 用 CV 看:Y 风险最大(CV=1.00),X 最有效率(CV=0.80)

结论: Y 看似安全(波动绝对值小),但单位回报承担的风险最大——中等回报但波动占比太高。

📌 面试/工作中的黄金提问:"波动是大了,但相比回报呢?"——这句话就能引出 CV。


九、总结:离散度工具箱

工具 公式 最佳使用场景
极差 Max − Min 快速粗略估计
MAD Σ|Xᵢ−X̄| / n 直观解释,非技术场合
总体方差 σ² Σ(Xᵢ−μ)² / N 已知总体数据
样本方差 s² Σ(Xᵢ−X̄)² / (n−1) 用样本推断总体
标准差 σ/s √方差 最常用的波动度量
CV s / X̄ 比较不同尺度的数据集
切比雪夫 1 − 1/k² 任意分布的安全估算

🔥 记住:没有人只看均值做决策。离散度是风险的语言,学会它,你就不再被数字忽悠。


十、测试题

题目 1

数据集:[2, 4, 4, 4, 5, 5, 7, 9],n = 8。

求极差(Range)。

A. 5 B. 6 C. 7 D. 8

题目 2

总体数据:[10, 12, 14, 16, 18]

求总体方差 σ²。

A. 8 B. 10 C. 6 D. 7

题目 3

接上题,求总体标准差 σ(保留两位小数)。

A. 2.83 B. 3.16 C. 2.45 D. 3.00

题目 4

以下关于样本方差分母使用 n−1 的说法,哪一个是最准确的?

A. 为了让方差变小一些,更保守 B. 因为丢掉了一个自由度,增加一点惩罚 C. 为了使样本方差成为总体方差的无偏估计量 D. 因为没有总体均值,只能用样本均值代替

题目 5

根据切比雪夫不等式,对于任何分布,至少有多大比例的数据落在均值 ± 2.5 个标准差范围内?

A. 75% B. 84% C. 88.9% D. 80%


十一、答案与解析

题目 1 答案:C. 7

解析: Range = X_max − X_min = 9 − 2 = 7

题目 2 答案:A. 8

解析: μ = (10+12+14+16+18) / 5 = 70/5 = 14

Xᵢ Xᵢ−μ (Xᵢ−μ)²
10 −4 16
12 −2 4
14 0 0
16 +2 4
18 +4 16

σ² = (16+4+0+4+16) / 5 = 40/5 = 8

题目 3 答案:A. 2.83

解析: σ = √8 ≈ 2.828 ≈ 2.83

题目 4 答案:C

解析:

C ✅ 无偏估计正是贝塞尔校正的数学目的——E(s²) = σ² D 部分正确但不完整——"没有总体均值"只是原因,不是答案 A 错误——除以 n−1 反而使方差变大 B 接近但不够精确——自由度解释是有道理的,但 CFA 标准答案是无偏估计

题目 5 答案:B. 84%

解析: 1 − 1/k² = 1 − 1/(2.5)² = 1 − 1/6.25 = 1 − 0.16 = 0.84 = 84%


十二、关键公式速记

公式 记忆口诀
Range = Max − Min "极差就是天花板减地板"
MAD = Σ|X−X̄|/n "绝对偏差取平均"
σ² = Σ(X−μ)²/N "总体方差:平方偏差除以总量"
s² = Σ(X−X̄)²/(n−1) "样本方差:N减一记心间"
σ = √σ² "标准差就是方差的平方根"
CV = s/X̄ "变异系数:标准差比均值"
切比雪夫 ≥ 1−1/k² "任何分布,保底估算"

CFA 一级 · L108 · 离散程度:极差、MAD、方差、标准差 · 中文版 · 2026-07-14

Topic: Don't Be Fooled by the Mean — Understanding Dispersion Is the Real Skill


1. Introduction: The Tale of Two Restaurants

Two restaurants with the same average customer rating of 4.0:

  • Restaurant A: Ratings [4, 4, 4, 4, 4] — every customer agrees
  • Restaurant B: Ratings [1, 1, 4, 5, 9] — some love it, some hate it

Same mean. Entirely different businesses.

Restaurant A is stable and predictable. Restaurant B is a roller coaster.

🔥 Core insight: The mean tells you "where the center is"; dispersion tells you "how spread out the data are." You must use both. Never one without the other.


2. What Is Dispersion?

2.1 Definition

Dispersion (Variability) measures how spread out or clustered data points are around their central tendency.

  • Low dispersion → data tightly clustered around the mean (low volatility)
  • High dispersion → data widely scattered (high volatility)

2.2 Relevance to Investing

Dispersion Implication Investment Takeaway
Low Stable returns, low fluctuations Suitable for conservative investors
High Returns swing widely High risk/high reward; requires higher risk tolerance

📌 Dispersion is the mathematical foundation of risk measurement in the CFA curriculum. From variance to beta to VaR — everything starts with these four tools.


3. The Four Dispersion Measures at a Glance

Measure Formula Idea Units Sensitivity to Outliers
Range Max − Min Original units 🔴 Extremely sensitive
MAD Average |Xᵢ − X̄| Original units 🟡 Moderately sensitive
Variance Average (Xᵢ − μ)² Squared units 🔴 Sensitive
Standard Deviation √Variance Original units 🟡 Moderately sensitive

4. Detailed Explanations

4.1 Range

Formula: Range = X_max − X_min

Advantage: Most intuitive — "the max minus the min."

Fatal flaw: Only uses two data points; ignores everything in between.

Practical Example:

Monthly returns for Fund A and Fund B: - A: [−2%, 1%, 2%, 3%, 5%] → Range = 5% − (−2%) = 7% - B: [−2%, −1%, 2%, 3%, 12%] → Range = 12% − (−2%) = 14%

Based on range alone, Fund B appears riskier. But Fund B has only one outlier (12%); all other months are similar to Fund A.

📌 The investment industry almost never uses range in isolation — a single outlier can completely distort it.

4.2 MAD — Mean Absolute Deviation

Formula:

$$\text{MAD} = \frac{\sum_{i=1}^{n} |X_i - \bar{X}|}{n}$$

Calculation steps: 1. Compute the mean X̄ 2. Compute the absolute deviation of each value from the mean: |Xᵢ − X̄| 3. Average all absolute deviations

Why absolute values?

Raw deviations (Xᵢ − X̄) always sum to zero (positives and negatives cancel out). Without absolute values, MAD would always be zero — meaningless.

Worked Example:

Dataset: [3, 5, 7, 9, 11]

X̄ = (3+5+7+9+11) / 5 = 7

Xᵢ Xᵢ − X̄ |Xᵢ − X̄|
3 −4 4
5 −2 2
7 0 0
9 +2 2
11 +4 4

MAD = (4+2+0+2+4) / 5 = 12/5 = 2.4

Interpretation: On average, each data point is 2.4 units away from the mean.

📌 MAD is easy to interpret and preserves original units, but absolute values are mathematically inconvenient (non-differentiable) — which is why variance and standard deviation dominate.

4.3 Variance — The Engine of Dispersion

4.3.1 Population Variance

Formula:

$$\sigma^2 = \frac{\sum_{i=1}^{N} (X_i - \mu)^2}{N}$$

4.3.2 Sample Variance

Formula:

$$s^2 = \frac{\sum_{i=1}^{n} (X_i - \bar{X})^2}{n-1}$$

4.3.3 Why Divide by n−1 Instead of n?

This is a must-know CFA Level 1 exam topic.

Intuition: The sample mean X̄ is computed from the sample itself, so it is "closer" to the sample data points than the true population mean μ would be. This makes deviations (Xᵢ − X̄) systematically smaller than (Xᵢ − μ).

Dividing by n−1 (instead of n) allocates the sum across fewer units — making the result larger and compensating for the downward bias.

Technical term: Bessel's Correction — makes the sample variance an unbiased estimator of the population variance.

🔑 Memory aid: "Population: N. Sample: n−1."

4.3.4 Variance Worked Example

Population data: [3, 5, 7, 9, 11], μ = 7

Xᵢ Xᵢ − μ (Xᵢ − μ)²
3 −4 16
5 −2 4
7 0 0
9 +2 4
11 +4 16

σ² = (16+4+0+4+16) / 5 = 40/5 = 8.0

4.3.5 The Core Flaw of Variance

Variance is expressed in squared units.

  • If data are "returns in %" → variance unit is "%²" (percent squared?!)
  • If data are "prices in dollars" → variance unit is "dollars²"

This is exactly why we need standard deviation — to bring us back to original units.

4.4 Standard Deviation — The "Final Form" of Dispersion

Formula:

Population standard deviation: $$\sigma = \sqrt{\sigma^2}$$

Sample standard deviation: $$s = \sqrt{s^2}$$

From the example above: σ = √8 = 2.828

Interpretation: Under a normal distribution, approximately 68% of observations fall within ±1 standard deviation of the mean.


5. Summary Comparison of the Four Measures

Dimension Range MAD Variance Std Dev
Uses all data? ❌ Only 2 points ✅ ✅ ✅
Units match original? ✅ ✅ ❌ Squared units ✅
Good mathematical properties? ❌ ⚠️ Absolute value non-differentiable ✅ Differentiable ✅ Differentiable
Robust to outliers? ❌ ⚠️ ❌ ⚠️
CFA exam importance ⭐ ⭐⭐ ⭐⭐⭐⭐⭐ ⭐⭐⭐⭐⭐

6. Chebyshev's Inequality

6.1 Definition

For any distribution, the proportion of observations falling within ± k standard deviations of the mean is at least:

$$1 - \frac{1}{k^2}, \quad k > 1$$

6.2 Application

k Minimum Proportion Calculation
1.25 36% 1 − 1/1.5625
1.5 55.6% 1 − 1/2.25
2 75% 1 − 1/4
2.5 84% 1 − 1/6.25
3 88.9% 1 − 1/9

📌 k=2 (≥75%) and k=3 (≥88.9%) are high-frequency CFA Level 1 exam points.

6.3 Comparison: Chebyshev vs. Normal Distribution

Interval Chebyshev (Any Distribution) Normal Distribution
μ ± 1σ Not applicable (k ≤ 1) ≈ 68%
μ ± 2σ At least 75% ≈ 95%
μ ± 3σ At least 89% ≈ 99.7%

🔥 Chebyshev is the "floor" estimate — safe for any distribution. The normal distribution gives precise proportions but only applies when normality holds.


7. Coefficient of Variation (CV)

7.1 Why Do We Need CV?

Consider: Fund A has mean return 12%, standard deviation 8%. Fund B has mean return 5%, standard deviation 4%. Which carries more risk per unit of return?

Standalone σ cannot answer this — the two funds have different means, and standard deviations scale accordingly.

7.2 Formula

$$\text{CV} = \frac{s}{\bar{X}} \quad \text{or} \quad \frac{\sigma}{\mu}$$

7.3 Worked Example

  • Fund A: CV_A = 8% / 12% = 0.667
  • Fund B: CV_B = 4% / 5% = 0.800

Fund B has the higher CV → more volatility per unit of return.

7.4 Limitation of CV

⚠️ CV breaks down when the mean is close to zero or negative (dividing by a near-zero value produces absurd results).


8. Practical Application: Portfolio Risk Identification

Scenario

You are analyzing annualized daily volatility for three stocks (standard deviations already computed):

Stock Mean Return Std Dev CV
X 15% 12% 0.80
Y 8% 8% 1.00
Z 20% 18% 0.90

Analysis: - By standard deviation alone: Z is riskiest (σ=18%), Y is safest (σ=8%) - By CV: Y is riskiest (CV=1.00), X is most efficient (CV=0.80)

Conclusion: Y seems safe in absolute terms, but per unit of return, Y carries the highest risk — modest returns but disproportionately high volatility.

📌 The golden question in interviews and at work: "The volatility is high, sure — but relative to the return?" — that question leads straight to CV.


9. Summary: Dispersion Toolkit

Tool Formula Best Use Case
Range Max − Min Quick rough estimate
MAD Σ|Xᵢ−X̄| / n Intuitive explanations, non-technical settings
Population Variance σ² Σ(Xᵢ−μ)² / N Known population data
Sample Variance s² Σ(Xᵢ−X̄)² / (n−1) Inferring population from sample
Standard Deviation σ/s √Variance Most widely used volatility measure
CV s / X̄ Comparing datasets of different scales
Chebyshev 1 − 1/k² Safe estimation for any distribution

🔥 Remember: No one makes decisions based on the mean alone. Dispersion is the language of risk. Learn it, and you will never be fooled by numbers.


10. Practice Questions

Question 1

Dataset: [2, 4, 4, 4, 5, 5, 7, 9], n = 8.

Calculate the range.

A. 5 B. 6 C. 7 D. 8

Question 2

Population data: [10, 12, 14, 16, 18]

Calculate the population variance σ².

A. 8 B. 10 C. 6 D. 7

Question 3

Using the same data, calculate the population standard deviation σ (to 2 decimal places).

A. 2.83 B. 3.16 C. 2.45 D. 3.00

Question 4

Regarding the use of n−1 as the denominator in sample variance, which of the following statements is most accurate?

A. It makes the variance smaller to be more conservative. B. One degree of freedom is lost, so a penalty is added. C. It makes the sample variance an unbiased estimator of the population variance. D. The population mean is unavailable, so the sample mean is used instead.

Question 5

According to Chebyshev's inequality, for any distribution, at least what proportion of observations falls within ±2.5 standard deviations of the mean?

A. 75% B. 84% C. 88.9% D. 80%


11. Answers and Explanations

Question 1 Answer: C. 7

Explanation: Range = X_max − X_min = 9 − 2 = 7

Question 2 Answer: A. 8

Explanation: μ = (10+12+14+16+18) / 5 = 70/5 = 14

Xᵢ Xᵢ−μ (Xᵢ−μ)²
10 −4 16
12 −2 4
14 0 0
16 +2 4
18 +4 16

σ² = (16+4+0+4+16) / 5 = 40/5 = 8

Question 3 Answer: A. 2.83

Explanation: σ = √8 ≈ 2.828 ≈ 2.83

Question 4 Answer: C

Explanation:

C ✅ Unbiased estimation is the precise mathematical purpose of Bessel's correction — E(s²) = σ². D is partially correct but incomplete — "no population mean" is the reason, but not the answer to "why n−1?" A is false — dividing by n−1 actually makes the variance larger, not smaller. B is close but imprecise — the degrees-of-freedom explanation has merit, but the CFA standard answer centers on unbiased estimation.

Question 5 Answer: B. 84%

Explanation: 1 − 1/k² = 1 − 1/(2.5)² = 1 − 1/6.25 = 1 − 0.16 = 0.84 = 84%


12. Key Formula Quick Reference

Formula Memory Aid
Range = Max − Min "Range is ceiling minus floor"
MAD = Σ|X−X̄|/n "Average of absolute deviations"
σ² = Σ(X−μ)²/N "Population variance: squared deviations over N"
s² = Σ(X−X̄)²/(n−1) "Sample variance: remember n−1"
σ = √σ² "Standard deviation is the square root of variance"
CV = s/X̄ "CV: standard deviation over mean"
Chebyshev ≥ 1−1/k² "Floor estimate for any distribution"

CFA Level 1 · L108 · Measures of Dispersion: Range, MAD, Variance, Standard Deviation · English Version · 2026-07-14

🔜 下一课 · L109

CFA Level 1 — L109:偏度(Skewness) — 一、什么是偏度? · 二、三种分布形态 · 三、均值、中位数、众数的位置关系