Standard II — Integrity of Capital Markets Module 1 · 15-20% Weight Lesson 107

📖 分位数:四分位数、百分位数

CFA Level 1 · L107 · Quantiles: Quartiles and Percentiles

课题:数据不会说话?分位数是你的翻译官


一、引言:为什么基金经理总能说"我们跑赢了一半同行"?

你见过这样的基金宣传吗: - "本基金业绩超过 75% 的同类产品" - "我们的回报率处于行业前 25%"

这些陈述背后的数学工具,就是分位数(Quantile)。

分位数是统计学中最实用、最直观的工具之一——它不关心平均值,只关心你在群体中的相对位置。

🔥 核心直觉:均值告诉你"大家平均赚多少";分位数告诉你"你比别人强多少"。在投资领域,后者往往更重要。


二、什么是分位数?

2.1 定义

分位数(Quantile) 是将一个有序数据集分割为等份的点。

给你一个已排序的数据集,第 p 分位数是这样一个值:有 p 比例的数据 ≤ 该值。

关键前提:数据必须先排序(从小到大)。 未排序的分位数没有意义。

2.2 通用公式

设数据量为 n,要求第 p 分位数:

位置指数:L_y = (n + 1) × y

其中 y 是分位点(百分位数时 y = p/100)。


三、三大核心分位数类型

3.1 四分位数(Quartiles)

将数据分成 4 等份。

四分位数 含义 分位点 别称
Q₁(第一四分位数) 25% 的数据 ≤ 此值 0.25 下四分位数
Q₂(第二四分位数) 50% 的数据 ≤ 此值 0.50 中位数
Q₃(第三四分位数) 75% 的数据 ≤ 此值 0.75 上四分位数

实战案例:基金评级

某平台将 1000 只基金按年化回报排序: - Q₁ = 第 250 名的回报:落后基金分界线 - Q₂ = 第 500 名的回报:市场平均水平 - Q₃ = 第 750 名的回报:优秀基金分界线

如果你的基金超过 Q₃ → 你在前 25%,可以上宣传材料了 📈

3.2 百分位数(Percentiles)

将数据分成 100 等份。

第 k 百分位数 = 有 k% 的数据 ≤ 该值。

关键对应关系: - P₂₅ = Q₁ - P₅₀ = Q₂ = 中位数 - P₇₅ = Q₃

实战案例:考试成绩

CFA 考试的通过线并非固定分数,而是基于最低及格分数(MPS)——本质上是一个经过校准的百分位数概念: - 你的得分如果超过 MPS 对应的百分位 → 通过 - 这是你和同期考生的相对竞赛,而非绝对分数

3.3 五分位数(Quintiles)与十分位数(Deciles)

类型 等份数 关键分位点 应用场景
五分位数 5 等份 20%, 40%, 60%, 80% 行业分组(如前 20%)
十分位数 10 等份 10%, 20%, ..., 90% 收入分配、风险分层

四、计算方法(核心考点)

4.1 位置公式

对于 n 个已排序数据,第 y 分位数(0 < y < 1)的位置:

L_y = (n + 1) × y

4.2 两步计算法

  1. 计算 L_y = (n+1) × y
  2. 如果 L_y 是整数 → 直接取第 L_y 个数据
  3. 如果 L_y 不是整数 → 线性插值

4.3 线性插值法

设 L_y = a + d,其中 a 是整数部分,d 是小数部分(0 < d < 1):

分位数 = X_a + d × (X_{a+1} − X_a)

4.4 完整计算示例

数据集(已排序): 3, 5, 7, 8, 9, 11, 13, 15

n = 8

求 Q₁(第 25 百分位): - L = (8+1) × 0.25 = 9 × 0.25 = 2.25 - a = 2, d = 0.25 - Q₁ = X₂ + 0.25 × (X₃ − X₂) = 5 + 0.25 × (7−5) = 5 + 0.5 = 5.5

求 Q₂(中位数 / 第 50 百分位): - L = (8+1) × 0.50 = 9 × 0.50 = 4.5 - a = 4, d = 0.5 - Q₂ = X₄ + 0.5 × (X₅ − X₄) = 8 + 0.5 × (9−8) = 8.5

求 Q₃(第 75 百分位): - L = (8+1) × 0.75 = 9 × 0.75 = 6.75 - a = 6, d = 0.75 - Q₃ = X₆ + 0.75 × (X₇ − X₆) = 11 + 0.75 × (13−11) = 11 + 1.5 = 12.5

4.5 四分位距(IQR)

IQR = Q₃ − Q₁

上例 IQR = 12.5 − 5.5 = 7.0

📌 IQR 是衡量数据中间 50% 离散程度的核心指标,不受极端值影响。


五、分位数的三大应用

5.1 箱线图(Box Plot)——可视化利器

      下边界       Q₁    Q₂    Q₃      上边界
       |          |─────|─────|          |
   o---+----------+=====+=====+----------+---o
  异常值          └── IQR ──┘          异常值

箱线图构造规则: - 箱体:Q₁ 到 Q₃ - 箱内线:Q₂(中位数) - 下须(Lower Whisker):Q₁ − 1.5 × IQR 范围内的最小值 - 上须(Upper Whisker):Q₃ + 1.5 × IQR 范围内的最大值 - 须外的点:异常值(Outliers)

5.2 异常值识别

温和异常值: 距离 Q₁ 或 Q₃ 超过 1.5 × IQR 极端异常值: 距离 Q₁ 或 Q₃ 超过 3.0 × IQR

📌 这是 CFA 一级常见考点:基于 IQR 的异常值判断 vs 基于标准差的判断。

5.3 相对位置评估

应用场景: 你想知道某只股票的表现到底处于什么水平。

  • 将同类 500 只股票排序
  • 你的股票 P₈₀ 分位 → 超过 80% 的同类
  • 这是一个比你报"平均涨了 15%"更有说服力的数字

六、注意事项与常见陷阱

6.1 陷阱一:忘记排序

❌ 错误做法:直接对原始数据取某个位置的值 ✅ 正确做法:必须先排序,再计算分位数

6.2 陷阱二:公式混淆

不同教材用的 (n+1) 系数可能略有差异(有些用 n,有些用 n-1),CFA 一级统一使用 L_y = (n+1) × y。

6.3 陷阱三:分位数 vs 分位点

  • 分位点(y): 一个比例值,如 0.25、0.75
  • 分位数(Quantile): 该分位点对应的数据值
  • 例:P₂₅(第 25 百分位数) = 5.5,其中 0.25 是分位点,5.5 是分位数

6.4 陷阱四:中位数的韧性

中位数不受极端值影响,均值会被极端值严重扭曲。

数据集 均值 中位数
[3, 5, 7, 8, 9, 11, 13, 15] 8.875 8.5
[3, 5, 7, 8, 9, 11, 13, 150] 25.75 8.5

均值从 8.875 暴涨到 25.75,中位数纹丝不动。 这就是为什么居民收入统计"平均数被平均"——要用中位数看真相。


七、总结:分位数工具箱备忘

工具 公式/定义 用途
位置指数 L_y = (n+1) × y 定位分位数值
Q₁ L_0.25 下四分位
Q₂ L_0.50 中位数
Q₃ L_0.75 上四分位
IQR Q₃ − Q₁ 中间 50% 离散度
异常值边界 Q₁ − 1.5×IQR / Q₃ + 1.5×IQR 识别异常值
线性插值 X_a + d × (X_{a+1} − X_a) 非整数位置计算

🔥 记住:统计学不是算数字,是讲故事。分位数帮你讲好"我在哪里"这个故事。


八、测试题

题目 1

数据集(已排序):[2, 4, 6, 8, 10, 12, 14, 16, 18, 20],n = 10。

求 Q₁。

A. 5.5 B. 5.25 C. 6.0 D. 5.75

题目 2

同上数据集,求第 80 百分位数 P₈₀。

A. 16.8 B. 16.0 C. 17.2 D. 17.8

题目 3

以下关于分位数的说法,哪一个是错误的?

A. 中位数是第 50 百分位数,也是第二四分位数 Q₂ B. 计算分位数前必须先对数据排序 C. IQR 对极端值非常敏感,因此不适合用于异常值检测 D. P₂₅ 和 Q₁ 是等价的


九、答案与解析

题目 1 答案:D. 5.75

解析:

Q₁ → y = 0.25 L = (10+1) × 0.25 = 11 × 0.25 = 2.75 a = 2, d = 0.75 Q₁ = X₂ + 0.75 × (X₃ − X₂) = 4 + 0.75 × (6 − 4) = 4 + 1.5 = 5.75

题目 2 答案:A. 16.8

解析:

P₈₀ → y = 0.80 L = (10+1) × 0.80 = 11 × 0.80 = 8.8 a = 8, d = 0.8 P₈₀ = X₈ + 0.8 × (X₉ − X₈) = 16 + 0.8 × (18−16) = 16 + 1.6 = 16.8

题目 3 答案:C

解析:

C 是错误的。IQR 的设计初衷正是不受极端值影响(因为它只看 Q₁ 和 Q₃,即中间 50% 数据)。基于 IQR 的异常值检测方法(箱线图法)是统计学中识别异常值的标准工具之一,正是因为 IQR 具有对极端值的稳健性。

A ✅ 中位数 = P₅₀ = Q₂ B ✅ 排序是计算分位数的前提 C ❌ 刚好说反——IQR 对极端值不敏感,天然适合异常值检测 D ✅ P₂₅ 和 Q₁ 都是 25% 分位点


十、关键公式速记

公式 记忆口诀
L_y = (n+1) × y "位置指数 = 总数加一乘分位"
IQR = Q₃ − Q₁ "箱体宽度就是四分位距"
插值 = X_a + d × (X_{a+1}−X_a) "基数加比例乘间距"
异常值下界 = Q₁ − 1.5×IQR "1.5 倍箱宽定异常"

CFA 一级 · L107 · 分位数:四分位数、百分位数 · 中文版 · 2026-07-13

Topic: Data Doesn't Speak — Quantiles Are Your Interpreter


1. Introduction: Why Every Fund Manager Claims "We Beat Half Our Peers"

You've seen this before: - "This fund outperforms 75% of peers" - "Our returns are in the top quartile"

Behind these statements lies a single mathematical tool: quantiles.

Quantiles are among the most practical and intuitive tools in statistics — they don't care about averages, only about your relative position in the group.

🔥 Core intuition: The mean tells you "how much everyone earns on average"; quantiles tell you "how much better you are than others." In investing, the latter often matters more.


2. What Is a Quantile?

2.1 Definition

A quantile is a value that divides a sorted dataset into equal-sized groups.

Given a sorted dataset, the p-th quantile is the value below which a proportion p of the data falls.

Critical prerequisite: data must be sorted (ascending) first. Unsorted quantiles are meaningless.

2.2 General Formula

For n data points, to find the y-th quantile:

Position index: L_y = (n + 1) × y

where y is the quantile point (for percentiles, y = k/100).


3. Three Core Quantile Types

3.1 Quartiles

Divide data into 4 equal parts.

Quartile Meaning Quantile Point Also Known As
Q₁ (First Quartile) 25% of data ≤ this value 0.25 Lower Quartile
Q₂ (Second Quartile) 50% of data ≤ this value 0.50 Median
Q₃ (Third Quartile) 75% of data ≤ this value 0.75 Upper Quartile

Real-World Example: Fund Ratings

A platform ranks 1,000 funds by annualized return: - Q₁ = return of the 250th fund: laggard cutoff - Q₂ = return of the 500th fund: market median - Q₃ = return of the 750th fund: top-performer cutoff

If your fund exceeds Q₃ → you're in the top 25%. Time for the marketing brochure 📈

3.2 Percentiles

Divide data into 100 equal parts.

The k-th percentile = k% of data ≤ this value.

Key Relationships: - P₂₅ = Q₁ - P₅₀ = Q₂ = Median - P₇₅ = Q₃

Real-World Example: Exam Scores

The CFA exam passing score is not a fixed number — it is based on the Minimum Passing Score (MPS), which is essentially a calibrated percentile concept: - If your score exceeds the MPS-aligned percentile → you pass - It's a relative competition against your cohort, not an absolute score threshold

3.3 Quintiles and Deciles

Type Equal Parts Key Points Application
Quintiles 5 20%, 40%, 60%, 80% Industry groupings (e.g., top 20%)
Deciles 10 10%, 20%, ..., 90% Income distribution, risk stratification

4. Calculation Method (Key Exam Topic)

4.1 Position Formula

For n sorted data points, the position of the y-th quantile (0 < y < 1):

L_y = (n + 1) × y

4.2 Two-Step Procedure

  1. Compute L_y = (n+1) × y
  2. If L_y is an integer → directly take the L_y-th data value
  3. If L_y is not an integer → linear interpolation

4.3 Linear Interpolation

Let L_y = a + d, where a is the integer part and d is the fractional part (0 < d < 1):

Quantile = X_a + d × (X_{a+1} − X_a)

4.4 Full Worked Example

Dataset (sorted): 3, 5, 7, 8, 9, 11, 13, 15

n = 8

Find Q₁ (25th percentile): - L = (8+1) × 0.25 = 9 × 0.25 = 2.25 - a = 2, d = 0.25 - Q₁ = X₂ + 0.25 × (X₃ − X₂) = 5 + 0.25 × (7−5) = 5 + 0.5 = 5.5

Find Q₂ (Median / 50th percentile): - L = (8+1) × 0.50 = 9 × 0.50 = 4.5 - a = 4, d = 0.5 - Q₂ = X₄ + 0.5 × (X₅ − X₄) = 8 + 0.5 × (9−8) = 8.5

Find Q₃ (75th percentile): - L = (8+1) × 0.75 = 9 × 0.75 = 6.75 - a = 6, d = 0.75 - Q₃ = X₆ + 0.75 × (X₇ − X₆) = 11 + 0.75 × (13−11) = 11 + 1.5 = 12.5

4.5 Interquartile Range (IQR)

IQR = Q₃ − Q₁

From the example: IQR = 12.5 − 5.5 = 7.0

📌 IQR is the core measure of dispersion for the middle 50% of data — it is unaffected by extreme values.


5. Three Key Applications of Quantiles

5.1 Box Plot — A Visualization Powerhouse

   Lower Fence    Q₁    Q₂    Q₃     Upper Fence
       |          |─────|─────|          |
   o---+----------+=====+=====+----------+---o
   Outlier              └── IQR ──┘       Outlier

Box Plot Construction Rules: - Box: Q₁ to Q₃ - Line inside box: Q₂ (median) - Lower Whisker: minimum value within Q₁ − 1.5 × IQR - Upper Whisker: maximum value within Q₃ + 1.5 × IQR - Points beyond whiskers: outliers

5.2 Outlier Detection

Mild outliers: more than 1.5 × IQR from Q₁ or Q₃ Extreme outliers: more than 3.0 × IQR from Q₁ or Q₃

📌 This is a common CFA Level 1 exam topic: IQR-based outlier detection vs. standard deviation-based detection.

5.3 Relative Position Assessment

Application: You want to know where a particular stock stands among its peers.

  • Sort 500 peer stocks by return
  • Your stock at the P₈₀ quantile → it beats 80% of peers
  • This is far more compelling than saying "the average return is 15%."

6. Pitfalls and Common Traps

6.1 Trap 1: Forgetting to Sort

❌ Wrong: taking a value at some position from raw data ✅ Correct: sort first, then compute quantiles

6.2 Trap 2: Formula Confusion

Different textbooks may use slightly different coefficients (some use n, some use n−1). CFA Level 1 uniformly uses L_y = (n+1) × y.

6.3 Trap 3: Quantile Point vs. Quantile Value

  • Quantile point (y): a proportion, e.g., 0.25, 0.75
  • Quantile value: the data value at that quantile point
  • Example: P₂₅ (25th percentile) = 5.5, where 0.25 is the quantile point, 5.5 is the quantile value

6.4 Trap 4: The Robustness of the Median

The median is unaffected by extreme values; the mean gets severely distorted.

Dataset Mean Median
[3, 5, 7, 8, 9, 11, 13, 15] 8.875 8.5
[3, 5, 7, 8, 9, 11, 13, 150] 25.75 8.5

The mean skyrockets from 8.875 to 25.75, while the median stays rock-solid. This is why income statistics often use the median to reveal the truth rather than the "average" that gets distorted by the ultra-rich.


7. Summary: Quantile Toolkit Cheat Sheet

Tool Formula/Definition Purpose
Position Index L_y = (n+1) × y Locate the quantile value
Q₁ L_0.25 Lower quartile
Q₂ L_0.50 Median
Q₃ L_0.75 Upper quartile
IQR Q₃ − Q₁ Middle 50% dispersion
Outlier Boundaries Q₁ − 1.5×IQR / Q₃ + 1.5×IQR Detect outliers
Linear Interpolation X_a + d × (X_{a+1} − X_a) Non-integer position calculation

🔥 Remember: statistics isn't about crunching numbers — it's about telling stories. Quantiles help you tell the story of "where I stand."


8. Practice Questions

Question 1

Dataset (sorted): [2, 4, 6, 8, 10, 12, 14, 16, 18, 20], n = 10.

Find Q₁.

A. 5.5 B. 5.25 C. 6.0 D. 5.75

Question 2

Using the same dataset, find the 80th percentile P₈₀.

A. 16.8 B. 16.0 C. 17.2 D. 17.8

Question 3

Which of the following statements about quantiles is incorrect?

A. The median is the 50th percentile and also the second quartile Q₂ B. Data must be sorted before computing quantiles C. IQR is highly sensitive to extreme values and therefore unsuitable for outlier detection D. P₂₅ and Q₁ are equivalent


9. Answers and Explanations

Question 1 Answer: D. 5.75

Explanation:

Q₁ → y = 0.25 L = (10+1) × 0.25 = 11 × 0.25 = 2.75 a = 2, d = 0.75 Q₁ = X₂ + 0.75 × (X₃ − X₂) = 4 + 0.75 × (6 − 4) = 4 + 1.5 = 5.75

Question 2 Answer: A. 16.8

Explanation:

P₈₀ → y = 0.80 L = (10+1) × 0.80 = 11 × 0.80 = 8.8 a = 8, d = 0.8 P₈₀ = X₈ + 0.8 × (X₉ − X₈) = 16 + 0.8 × (18−16) = 16 + 1.6 = 16.8

Question 3 Answer: C

Explanation:

C is incorrect. The IQR is specifically designed to be insensitive to extreme values (because it only considers Q₁ and Q₃ — the middle 50% of data). IQR-based outlier detection (the box plot method) is one of the standard tools in statistics for identifying outliers, precisely because of IQR's robustness to extreme values.

A ✅ Median = P₅₀ = Q₂ B ✅ Sorting is a prerequisite for quantile calculation C ❌ Gets it exactly backwards — IQR is insensitive to extremes, making it naturally suited for outlier detection D ✅ P₂₅ and Q₁ are both the 25% quantile point


10. Key Formula Quick Reference

Formula Memory Aid
L_y = (n+1) × y "Position = sample size plus one times quantile"
IQR = Q₃ − Q₁ "Box width equals interquartile range"
Interpolation = X_a + d × (X_{a+1}−X_a) "Base plus fraction times gap"
Outlier Lower Bound = Q₁ − 1.5×IQR "1.5 box-widths defines outliers"

CFA Level 1 · L107 · Quantiles: Quartiles and Percentiles · English Version · 2026-07-13

🔜 下一课 · L108

CFA 一级 · L108 · 离散程度:极差、MAD、方差、标准差 — 课题:别被平均值骗了——学会看"离散度"才是真本事 · 一、引言:两碗汤的故事 · 二、什么是离散程度(Dispersion)?