Standard II — Integrity of Capital Markets Module 1 · 15-20% Weight Lesson 125

📖 抽样方法:简单随机、分层、系统

CFA Level 1 · L125 · Sampling Methods: Simple Random, Stratified, Systematic

定量方法(Quantitative Methods)— 抽样与估计 · 第 1 课


一、为什么需要抽样?

1.1 总体 vs 样本

概念 英文 定义 示例
总体 Population 研究对象的全部个体集合 深交所全部 2800+ 只股票
样本 Sample 从总体中抽取的一部分个体 从中选取 200 只股票
参数 Parameter 描述总体特征的数值(通常未知) 全部股票的平均市盈率 μ
统计量 Statistic 从样本计算出的数值(用于估计参数) 200 只股票的平均市盈率 $\bar{x}$

💡 核心逻辑: 因为调查整个总体成本太高 / 不可能(如破坏性测试),我们用样本统计量来估计总体参数,这就是统计推断(Statistical Inference)的基础。

1.2 CFA 考试中的常见情境

  • 基金经理想了解某板块整体盈利能力 → 不可能分析所有公司 → 抽样分析
  • 品控部门测试产品寿命 → 不能把所有产品都用坏 → 抽样测试
  • 分析师调查消费者信心 → 不可能采访全国所有人 → 抽样调查

二、抽样方法的分类总览

抽样方法(Sampling Methods)
│
├── 概率抽样(Probability Sampling)
│   ├── 简单随机抽样(Simple Random Sampling)
│   ├── 分层随机抽样(Stratified Random Sampling)
│   └── 系统抽样(Systematic Sampling)
│
└── 非概率抽样(Non-Probability Sampling)
    ├── 便利抽样(Convenience Sampling)
    ├── 判断抽样(Judgmental Sampling)
    └── 配额抽样(Quota Sampling)

📌 CFA 一级重点考概率抽样的三种方法及其优缺点,非概率抽样知道定义和局限性即可。


三、简单随机抽样(Simple Random Sampling)

3.1 定义

从总体中抽取样本,使得每个个体被抽中的概率完全相等,且每次抽取相互独立。

3.2 操作方式

  • 给总体中每个个体编号(如 1 到 N)
  • 使用随机数生成器(Excel RAND()、Python random.sample())抽取 n 个号码
  • 对应号码的个体进入样本

3.3 实战案例

案例: 分析师要从标普 500 成分股中随机选取 50 只做财务比率分析。

  • 给 500 只股票编号 1-500
  • 用 =RANDBETWEEN(1,500) 生成 50 个不重复的随机数
  • 被选中的 50 只组成样本

每只股票被选中的概率 = 50/500 = 10%

3.4 优点

优点 说明
无偏性 样本均值的期望 = 总体均值,$E(\bar{x}) = \mu$
简单直观 计算抽样误差最简单
理论基础成熟 中心极限定理的基础假设就是简单随机抽样

3.5 缺点

缺点 说明
需要完整名单 必须有总体所有个体的列表(sampling frame)
成本高 当总体分散在各地时,实施困难
可能不具代表性 纯随机可能导致某些子群体被遗漏(样本容量小时)
不适合异质性强的总体 如果总体内差异很大,需极大样本量才能准确

四、分层随机抽样(Stratified Random Sampling)

4.1 定义

先将总体按某个特征(分层变量)分成若干互不重叠的层(Strata),然后在每一层内部独立进行简单随机抽样。

4.2 操作流程

步骤 1:选择分层变量
        ↓
步骤 2:将总体划分为 L 个互斥的层(H₁, H₂, ..., Hₗ)
        ↓
步骤 3:在每个层内独立进行简单随机抽样
        ↓
步骤 4:合并各层样本,得到最终样本

4.3 分层变量的选择标准

  • 与研究对象高度相关的特征
  • 层内个体特征相似(同质性高)
  • 层间个体差异明显(异质性高)

4.4 实战案例

案例: 研究中国 A 股市场的估值水平。总体为全部 A 股。

  • 分层变量: 行业板块(金融、科技、消费、医药、能源等)
  • 原因: 不同行业的 PE 倍数水平差异巨大(科技 PE 普遍 > 金融 PE)
  • 做法: 在每个行业内各随机抽取 30 只股票
  • 效果: 保证每个行业在样本中都有代表,避免"全是科技股"的偏误

4.5 两种分配方式

方式 英文 规则 适用场景
等比例分配 Proportional Allocation 每层样本量 ∝ 层在总体中的占比 各层方差相近
非等比例分配 Disproportional Allocation 按分析需求调整各层样本量 某些层方差大,需更多样本来提高精度

4.6 优点

优点 说明
代表性更强 确保每个子群体都被覆盖
精度更高 层内同质 → 层内方差小 → 整体估计更精确
可做子群体分析 每个层可单独分析

4.7 缺点

缺点 说明
需要分层信息 必须知道总体的分层结构
设计复杂 选择分层变量、决定样本分配都需要专业知识
分层变量选择错误 如果分层变量与研究对象无关 → 徒增复杂度,不提升精度

五、系统抽样(Systematic Sampling)

5.1 定义

将总体按某种顺序排列后,每隔固定间隔 k 抽取一个个体。k 称为抽样间隔(Sampling Interval)。

5.2 操作流程

$$k = \frac{N}{n}$$

  • 随机选择一个起点(在 1 到 k 之间)
  • 每隔 k 个抽取一个
  • 抽取的第 i 个个体编号 = 起点 + (i−1) × k

5.3 实战案例

案例: 某工厂生产线每天生产 10,000 件产品,质检员要抽取 200 件检测。

  • k = 10,000 / 200 = 50
  • 在 1-50 之间随机选起点,假设为 23
  • 抽取第 23、73、123、173、... 件产品

这就是等距抽样的应用。

5.4 优点

优点 说明
操作简单 不需要随机数表,只需确定间隔
分布均匀 样本自动覆盖总体"全范围"
不需要完整的名单 只需知道总体大小 N 和排列顺序

5.5 缺点

缺点 说明
⚠️ 周期性偏误 如果总体排列存在周期性模式,且 k 恰好与周期同步,会产生严重偏误
方差的估计复杂 系统样本不是严格独立的,标准误差公式不同

5.6 ⚠️ 周期性偏误——CFA 高频考点

经典例子: 分析师要抽样评估某零售商的日均销售额。每天取一个时间点:

  • 抽样间隔 k = 7 天
  • 如果每天都固定选周六 → 周六销售额异常高 → 严重高估日均销售额
  • 这就是周期性偏误(Periodicity Bias)

应对方法: 检查总体是否存在周期性模式;如有,改用简单随机抽样。


六、三种概率抽样方法对比

特征 简单随机抽样 分层随机抽样 系统抽样
核心原理 每个个体等概率被抽中 先分层,层内随机抽 等间距抽取
需要总体名单? ✅ 必须 ✅ 必须(每层) ❌ 不需要完整名单
代表性 一般(样本小可能偏) ⭐ 最强 较好
操作难度 中 高 低
主要风险 抽样框架偏误 分层变量选错 周期性偏误
CFA 考试权重 ⭐⭐⭐ ⭐⭐⭐ ⭐⭐

七、非概率抽样(了解即可)

CFA 一级要求知道非概率抽样不能用于统计推断,因为无法计算抽样误差。

方法 定义 致命缺陷
便利抽样 选最容易获取的个体(街头发问卷) 选择性偏误严重
判断抽样 凭专家判断选择样本 主观性强,不可复制
配额抽样 先设各类别的配额,再便利选取 配额内仍不是随机选

八、关键概念辨析

8.1 抽样误差 vs 抽样偏误

抽样误差(Sampling Error) 抽样偏误(Sampling Bias)
定义 样本统计量与总体参数之间的随机差异 系统性高估或低估
是否可消除? ❌ 永远存在,但可通过增大样本量减小 ✅ 理论上可通过改善抽样设计消除
来源 样本只是总体的一部分,天然有波动 抽样方法设计不当
CFA 考点 增大 n 可减少,但不能消除 概率抽样可避免偏误

8.2 抽样框架(Sampling Frame)

定义: 从中抽取样本的"名单"或"列表"。

关键风险: 抽样框架 ≠ 总体 → 抽样框架偏误(Sampling Frame Bias)

例子: 用电话簿做抽样框架调查居民收入 → 漏掉了没有座机的人(可能收入偏低)→ 高估收入


九、CFA 考试答题思路

遇到抽样方法选择题时,按以下流程判断:

【判读题】
  ├── 题目是否描述了"先分组再抽样"?
  │   ├── 是 → 分层随机抽样(关键:层内随机)
  │   └── 否 → 继续
  ├── 题目是否描述了"每隔 k 个抽一个"?
  │   ├── 是 → 系统抽样(⚠️ 检查是否有周期性风险)
  │   └── 否 → 继续
  └── 题目是否描述了"完全随机,无任何预处理"?
      └── 是 → 简单随机抽样

十、测试题


题 1

分析师要将某市 50 所高中分为"重点高中"和"普通高中"两类,然后在每类中各随机抽取 5 所学校进行调查。这种抽样方法是:

A. 简单随机抽样 B. 分层随机抽样 C. 系统抽样 D. 判断抽样


题 2

某超市要从 20,000 名会员中抽取 500 名做满意度调查。按会员 ID 排序后,每隔 40 名抽取一位。这种方法属于:

A. 简单随机抽样 B. 分层随机抽样 C. 系统抽样 D. 配额抽样


题 3(易错题)

关于分层随机抽样,以下哪项说法是正确的?

A. 分层抽样一定比简单随机抽样更精确 B. 分层抽样的层内个体应该差异尽可能大 C. 分层变量必须与研究变量高度相关才能提升精度 D. 分层抽样中各层必须使用相同的样本量


题 4

以下哪种抽样方法必须拥有总体的完整名单?

A. 简单随机抽样 B. 系统抽样 C. 便利抽样 D. 判断抽样


题 5(情景题)

某分析师想了解纽交所全部上市公司的分红政策。他把所有股票按代码字母顺序排列,选代码以"A"开头的全部公司作为样本。该样本最可能的问题是:

A. 抽样误差过大 B. 周期性偏误 C. 抽样框架偏误 D. 便利抽样导致的系统性偏误


十一、答案与解析


【题 1 答案】B — 分层随机抽样

先按"重点/普通"分类(分层),再在每层内随机抽 → 这就是分层随机抽样的标准操作。

🧠 判断分层的关键词:"先分成……组/类/层"+"各类内随机抽取"


【题 2 答案】C — 系统抽样

等间隔抽取(每隔 40 名取一个)→ 系统抽样的定义性特征。

🧠 系统抽样的关键词:"每隔……"、"等距"、"第 k 个"。注意会员 ID 排序不存在周期性风险(ID 是随机分配的序号)。


【题 3 答案】C

选项 判断 理由
A ❌ 分层变量选错 → 分层抽样不一定更精确
B ❌ 层内应该同质(差异小),层间应该异质(差异大)
C ✅ 分层变量与研究变量相关,"层间差异大、层内差异小"才能提升精度
D ❌ 非等比例分配(disproportional)也是允许的

🧠 分层抽样的效率来源于:层内同质 → 层内方差小 → 总标准误差小。如果分层变量与研究变量无关,则无法实现这个效果。


【题 4 答案】A — 简单随机抽样

简单随机抽样要求给每个个体赋予相同的被抽中概率,这必须基于完整名单。

  • 系统抽样:只需知道总人数和排序,不需要一一对应的完整名单
  • 便利 / 判断抽样:非概率抽样,不要求名单

【题 5 答案】D — 便利抽样导致的系统性偏误

选代码以"A"开头的公司 → 这是便利抽样(因为容易找到),不是随机抽样。

  • 字母顺序可能与公司特征相关(如中国公司在美股多用特定字母)
  • 这个样本不能代表整个纽交所
  • 存在严重的选择性偏误

🧠 这个题考的是非概率抽样的局限性——不能用于统计推断。代码以"A"开头 ≠ 随机。


十二、本课小结

要点 一句话总结
抽样目的 用样本统计量估计总体参数
简单随机 最基础、最无偏,但需要完整名单
分层随机 先分组再抽样,层内同质、层间异质 → 精度最高
系统抽样 等间距抽取,操作最简单,但警惕周期性偏误
非概率抽样 不能计算抽样误差 → 不能做统计推断

📚 下一课 L126:抽样误差与标准误——抽样误差从哪来?标准误怎么算?中心极限定理如何保证"样本均值 ≈ 正态"?

Quantitative Methods — Sampling and Estimation · Lesson 1


1. Why Sampling?

1.1 Population vs Sample

Concept Definition Example
Population The entire collection of individuals under study All 2,800+ stocks listed on the Shenzhen Stock Exchange
Sample A subset of individuals drawn from the population 200 stocks selected from the exchange
Parameter A numerical characteristic of the population (usually unknown) The true mean P/E ratio μ of all stocks
Statistic A numerical value computed from the sample (used to estimate the parameter) The sample mean P/E ratio $\bar{x}$ of 200 stocks

💡 Core Logic: Because surveying an entire population is too costly or impossible (e.g., destructive testing), we use sample statistics to estimate population parameters. This forms the foundation of Statistical Inference.

1.2 Common CFA Scenarios

  • A fund manager wants to assess the overall profitability of a sector → cannot analyze every company → takes a sample
  • Quality control tests product lifespan → cannot destroy every unit → samples for testing
  • An analyst surveys consumer confidence → cannot interview the entire nation → uses sampling

2. Classification of Sampling Methods

Sampling Methods
│
├── Probability Sampling
│   ├── Simple Random Sampling
│   ├── Stratified Random Sampling
│   └── Systematic Sampling
│
└── Non-Probability Sampling
    ├── Convenience Sampling
    ├── Judgmental Sampling
    └── Quota Sampling

📌 CFA Level 1 focuses heavily on the three probability sampling methods and their pros/cons. Non-probability sampling only requires knowing the definitions and limitations.


3. Simple Random Sampling

3.1 Definition

Drawing a sample from a population such that every individual has an equal probability of being selected, and each draw is independent of others.

3.2 Procedure

  • Assign a unique number to each individual in the population (1 to N)
  • Use a random number generator (Excel RAND(), Python random.sample()) to select n numbers
  • The individuals corresponding to those numbers form the sample

3.3 Practical Example

Case: An analyst wants to randomly select 50 stocks from the S&P 500 for financial ratio analysis.

  • Number all 500 stocks from 1 to 500
  • Use =RANDBETWEEN(1,500) to generate 50 unique random numbers
  • The selected 50 stocks form the sample

Probability of any individual stock being selected = 50/500 = 10%

3.4 Advantages

Advantage Explanation
Unbiased Expected value of the sample mean equals the population mean: $E(\bar{x}) = \mu$
Simple & intuitive Easiest method for computing sampling error
Strong theoretical foundation The Central Limit Theorem is based on simple random sampling assumptions

3.5 Disadvantages

Disadvantage Explanation
Requires a complete list Must have a sampling frame listing all population members
High cost Difficult to implement when population is geographically dispersed
May lack representativeness Pure randomness may miss certain subgroups (especially with small samples)
Poor for heterogeneous populations Large sample sizes needed when within-population variance is high

4. Stratified Random Sampling

4.1 Definition

First divide the population into mutually exclusive strata based on a characteristic (stratification variable), then perform independent simple random sampling within each stratum.

4.2 Procedure

Step 1: Choose a stratification variable
        ↓
Step 2: Divide the population into L mutually exclusive strata (H₁, H₂, ..., Hₗ)
        ↓
Step 3: Perform simple random sampling independently within each stratum
        ↓
Step 4: Combine all stratum samples into the final sample

4.3 Criteria for Choosing Stratification Variables

  • Must be highly correlated with the variable under study
  • Individuals within each stratum should be homogeneous (similar)
  • Individuals across different strata should be heterogeneous (different)

4.4 Practical Example

Case: Studying valuation levels in China's A-share market. Population = all A-share stocks.

  • Stratification variable: Industry sector (financials, technology, consumer, healthcare, energy, etc.)
  • Reason: P/E ratios differ dramatically across sectors (tech P/E >> financials P/E)
  • Approach: Randomly select 30 stocks from each sector
  • Effect: Ensures every sector is represented, avoiding bias such as "all technology stocks"

4.5 Two Allocation Methods

Method Rule Best Used When
Proportional Allocation Sample size per stratum ∝ stratum's share of population Variance is similar across strata
Disproportional Allocation Adjust sample sizes per stratum based on analytical needs Some strata have larger variance, requiring more observations to improve precision

4.6 Advantages

Advantage Explanation
Stronger representativeness Ensures every subgroup is covered
Higher precision Within-stratum homogeneity → lower within-stratum variance → more precise overall estimates
Enables subgroup analysis Each stratum can be analyzed separately

4.7 Disadvantages

Disadvantage Explanation
Requires stratification information Must know the population's stratification structure
Complex design Choosing variables and allocating samples requires expertise
Wrong stratification variable If uncorrelated with the study variable → adds complexity without improving precision

5. Systematic Sampling

5.1 Definition

After arranging the population in some order, select every k-th individual at a fixed interval. k is called the sampling interval.

5.2 Procedure

$$k = \frac{N}{n}$$

  • Randomly select a starting point (between 1 and k)
  • Select every k-th individual thereafter
  • The i-th selected individual number = starting point + (i−1) × k

5.3 Practical Example

Case: A factory produces 10,000 units per day. A quality inspector needs to sample 200 units for testing.

  • k = 10,000 / 200 = 50
  • Randomly select a starting point between 1 and 50, say 23
  • Select units #23, #73, #123, #173, ...

This is a classic application of interval sampling.

5.4 Advantages

Advantage Explanation
Simple to execute No random number table needed — just determine the interval
Even coverage Sample automatically covers the "full range" of the population
No complete list required Only need population size N and a sequential ordering

5.5 Disadvantages

Disadvantage Explanation
⚠️ Periodicity bias If the population has a cyclical pattern and k aligns with the cycle, severe bias results
Complex variance estimation Systematic samples are not strictly independent; standard error formulas differ

5.6 ⚠️ Periodicity Bias — High-Frequency CFA Topic

Classic example: An analyst samples a retailer's daily sales. Takes one data point per day:

  • Sampling interval k = 7 days
  • If the selected day always falls on Saturday → Saturday sales are abnormally high → severely overestimates average daily sales
  • This is Periodicity Bias

Mitigation: Check whether the population exhibits a cyclical pattern; if so, switch to simple random sampling.


6. Comparison of the Three Probability Sampling Methods

Feature Simple Random Stratified Random Systematic
Core principle Every individual has equal selection probability Divide into strata first, then random-sample within each Select at fixed intervals
Requires complete population list? ✅ Required ✅ Required (per stratum) ❌ Not required
Representativeness Moderate (may be biased with small samples) ⭐ Strongest Good
Operational difficulty Medium High Low
Main risk Sampling frame bias Wrong stratification variable Periodicity bias
CFA exam weight ⭐⭐⭐ ⭐⭐⭐ ⭐⭐

7. Non-Probability Sampling (Overview)

CFA Level 1 requires knowing that non-probability sampling cannot be used for statistical inference because sampling error cannot be calculated.

Method Definition Fatal Flaw
Convenience Sampling Select the most easily accessible individuals (e.g., street surveys) Severe selection bias
Judgmental Sampling Select sample based on expert judgment Highly subjective, not replicable
Quota Sampling Set quotas for each category first, then use convenience selection within Selection within quotas remains non-random

8. Key Conceptual Distinctions

8.1 Sampling Error vs Sampling Bias

Sampling Error Sampling Bias
Definition Random difference between sample statistic and population parameter Systematic overestimation or underestimation
Can it be eliminated? ❌ Always exists, but can be reduced by increasing sample size ✅ Theoretically eliminable by improving sampling design
Source Sample is only part of the population — natural variability Flawed sampling method design
CFA takeaway Increasing n reduces but cannot eliminate it Probability sampling avoids bias

8.2 Sampling Frame

Definition: The "list" or "directory" from which the sample is drawn.

Key Risk: Sampling frame ≠ Population → Sampling Frame Bias

Example: Using a telephone directory as the sampling frame to survey household income → misses people without landlines (who may have lower incomes) → overestimates income


9. CFA Exam Decision Tree

When faced with a sampling method multiple-choice question, follow this logic:

【Diagnostic Flow】
  ├── Does the question describe "grouping first, then sampling"?
  │   ├── Yes → Stratified Random Sampling (key: random within strata)
  │   └── No → Continue
  ├── Does the question describe "every k-th individual"?
  │   ├── Yes → Systematic Sampling (⚠️ check for periodicity risk)
  │   └── No → Continue
  └── Does the question describe "completely random, no pre-processing"?
      └── Yes → Simple Random Sampling

10. Practice Questions


Question 1

An analyst divides 50 high schools in a city into "key schools" and "regular schools," then randomly selects 5 schools from each category for a survey. This sampling method is:

A. Simple random sampling B. Stratified random sampling C. Systematic sampling D. Judgmental sampling


Question 2

A supermarket wants to survey 500 out of its 20,000 members. After sorting by member ID, every 40th member is selected. This method is:

A. Simple random sampling B. Stratified random sampling C. Systematic sampling D. Quota sampling


Question 3 (Tricky)

Regarding stratified random sampling, which of the following statements is correct?

A. Stratified sampling is always more precise than simple random sampling B. Individuals within a stratum should be as different as possible C. The stratification variable must be highly correlated with the study variable to improve precision D. Every stratum must use the same sample size in stratified sampling


Question 4

Which of the following sampling methods must have a complete list of the population?

A. Simple random sampling B. Systematic sampling C. Convenience sampling D. Judgmental sampling


Question 5 (Scenario)

An analyst wants to study the dividend policies of all NYSE-listed companies. She sorts all stocks alphabetically by ticker and selects all companies whose tickers start with "A" as the sample. The most likely problem with this sample is:

A. Excessive sampling error B. Periodicity bias C. Sampling frame bias D. Systematic bias from convenience sampling


11. Answers and Explanations


【Question 1 Answer】B — Stratified Random Sampling

First classify into "key/regular" (stratification), then random-sample within each category → this is the standard procedure for stratified random sampling.

🧠 Keywords for identifying stratified sampling: "first divide into... groups/categories/strata" + "randomly select within each"


【Question 2 Answer】C — Systematic Sampling

Selecting at a fixed interval (every 40th member) → the defining characteristic of systematic sampling.

🧠 Keywords for systematic sampling: "every..." "fixed interval" "k-th". Note: sorting by member ID presents no periodicity risk since IDs are randomly assigned.


【Question 3 Answer】C

Option Verdict Rationale
A ❌ If the stratification variable is poorly chosen, stratified sampling may not be more precise
B ❌ Within-stratum should be homogeneous (small differences); between-strata should be heterogeneous (large differences)
C ✅ When the stratification variable is correlated with the study variable, "large between-strata variance, small within-stratum variance" improves precision
D ❌ Disproportional allocation is also permissible

🧠 The efficiency of stratified sampling comes from: within-stratum homogeneity → low within-stratum variance → low overall standard error. If the stratification variable is unrelated to the study variable, this benefit does not materialize.


【Question 4 Answer】A — Simple Random Sampling

Simple random sampling requires giving every individual an equal probability of selection, which can only be ensured with a complete list.

  • Systematic sampling: only requires knowing population size N and the ordering, not a one-to-one complete list
  • Convenience / Judgmental sampling: non-probability methods, no list required

【Question 5 Answer】D — Systematic bias from convenience sampling

Selecting all companies whose tickers start with "A" → this is convenience sampling (easy to find), not random sampling.

  • Alphabetical order may correlate with company characteristics (e.g., Chinese companies listed in the US often cluster around certain letters)
  • This sample cannot represent the entire NYSE
  • Severe selection bias exists

🧠 This question tests the limitation of non-probability sampling — it cannot be used for statistical inference. "Ticker starts with A" ≠ random.


12. Lesson Summary

Key Point One-Liner
Purpose of sampling Use sample statistics to estimate population parameters
Simple random Most basic, most unbiased, but requires a complete list
Stratified random Group first, then sample — within-strata homogeneous, between-strata heterogeneous → highest precision
Systematic Fixed-interval selection — easiest to execute, but watch for periodicity bias
Non-probability sampling Cannot compute sampling error → cannot be used for statistical inference

📚 Next Lesson L126: Sampling Error and Standard Error — Where does sampling error come from? How is standard error calculated? How does the Central Limit Theorem guarantee that "the sample mean ≈ normal"?

🔜 下一课 · L126

CFA 一级 · L126 · 中心极限定理 — 一、从一个问题出发 · 二、中心极限定理:核心陈述 · 三、直观理解:骰子实验