定量方法(Quantitative Methods)— 抽样与估计 · 第 2 课
一、从一个问题出发
假设你想估计深交所全部股票的平均市盈率(PE)。总体有 2800+ 只股票,你随机抽取了 100 只,算出一个样本均值 $\bar{x}$ = 15.3。
问题: 这个 15.3 离真实的总体均值 μ 有多远?如果重复抽 1000 次(每次抽 100 只),这些 $\bar{x}$ 会呈现什么分布?
这个问题的答案,就是中心极限定理(Central Limit Theorem,简称 CLT)。
二、中心极限定理:核心陈述
2.1 定义
中心极限定理(CLT): 对于任意一个均值为 μ、方差为 σ² 的总体,从中抽取容量为 n 的简单随机样本,当样本容量 n 足够大时,样本均值 $\bar{x}$ 的抽样分布近似服从正态分布,其均值为 μ,方差为 σ²/n。
用数学语言表达:
$$\bar{x} \sim N\left(\mu, \frac{\sigma^2}{n}\right) \quad \text{(当 n 足够大时)}$$
2.2 CLT 的三个核心结论
| # | 结论 | 含义 |
|---|---|---|
| 1 | 样本均值的期望 = 总体均值 | $E(\bar{x}) = \mu$ — 无偏性 |
| 2 | 样本均值的方差 = σ²/n | $Var(\bar{x}) = \frac{\sigma^2}{n}$ — 样本越大,波动越小 |
| 3 | 分布形态趋近正态 | 无论总体是什么分布,只要 n 够大,$\bar{x}$ 的分布都趋向正态 |
🔥 第三条是 CLT 的精髓: 总体可以是均匀分布、指数分布、甚至高度偏态——只要 n 够大,样本均值的分布就近似正态。这就是 CLT 如此强大的原因。
三、直观理解:骰子实验
3.1 单个骰子的分布
掷一颗公平骰子,结果 X 的分布是离散均匀分布(1-6 各 1/6):
P(X)
│
│ █ █ █ █ █ █
│ █ █ █ █ █ █
└──┴──┴──┴──┴──┴──
1 2 3 4 5 6 → 完全不是正态!
- μ = 3.5,σ² ≈ 2.92
3.2 掷两颗骰子,取平均值
掷 2 颗骰子取平均值 $\bar{x}_2$:
P(平均)
│ █
│ █████
│ █████████
│ █████████████
└────────────────
1 2 3 4 5 6 → 开始出现三角形状
3.3 掷 30 颗骰子,取平均值
掷 30 颗骰子取平均值 $\bar{x}_{30}$:
P(平均)
│ ╭──╮
│ ╱ ╲
│ ╱ ╲
│ ╱ ╲
└────────────────────
2.5 3.0 3.5 4.0 4.5 → 近似正态!
💡 骰子的总体分布是完全非正态的(均匀离散),但当 n=30 时,样本均值的分布已经非常接近正态了。
四、CLT 的关键条件
4.1 "n 足够大"到底多大?
| 总体分布特征 | 所需最小样本量 | 原因 |
|---|---|---|
| 总体本身接近正态 | n ≥ 10 即可 | 正态总体的抽样分布本身就是正态 |
| 总体对称但非正态(如均匀分布) | n ≥ 15~20 | 对称性有助于快速收敛 |
| 总体偏态(如收入分布、股票收益) | n ≥ 30 | CFA 常规标准 |
| 总体极度偏态(如保险理赔额) | n 可能需要更大 | 需要更多样本对抗偏态 |
📌 CFA 一级考试标准答案:n ≥ 30 即可认为样本量"足够大"。这是约定俗成的经验法则(Rule of Thumb)。
4.2 必须满足的前提
| 条件 | 说明 |
|---|---|
| 简单随机样本 | 每个个体被抽中概率相等,抽取独立 |
| 总体方差 σ² 有限 | 总体方差必须存在且有限 |
| n 足够大 | 通常取 n ≥ 30 |
五、CLT 的数学内涵
5.1 抽样分布(Sampling Distribution)
定义: 从同一总体中重复抽取容量相同的多个样本,所有样本均值的分布就是"$\bar{x}$ 的抽样分布"。
5.2 三条性质
$$\text{样本均值的抽样分布:}$$
| 参数 | 公式 | 说明 |
|---|---|---|
| 均值 | $E(\bar{x}) = \mu$ | $\bar{x}$ 是 μ 的无偏估计量 |
| 方差 | $Var(\bar{x}) = \frac{\sigma^2}{n}$ | 样本量增大,方差减小 |
| 标准差(标准误) | $SE(\bar{x}) = \frac{\sigma}{\sqrt{n}}$ | 标准误与 √n 成反比 |
5.3 标准误的含义
标准误(Standard Error) 是样本均值抽样分布的标准差,衡量 $\bar{x}$ 围绕 μ 的波动程度。
$$\text{精度} \propto \frac{1}{\sqrt{n}}$$
| 样本量 n | 标准误(相对于 σ) | 精度提升 |
|---|---|---|
| 10 | σ/√10 ≈ 0.316σ | — |
| 30 | σ/√30 ≈ 0.183σ | 比 n=10 提升 73% |
| 100 | σ/√100 = 0.1σ | 比 n=30 提升 83% |
| 400 | σ/√400 = 0.05σ | 比 n=100 提升 100% |
🔑 关键规律: 要把精度提高一倍(标准误减半),需要将样本量增加到原来的 4 倍。精度提升代价呈平方增长。
六、实战案例
6.1 基金收益分析
场景: 某对冲基金声称其策略的月收益率均值为 1.2%,标准差为 3.5%。监管机构随机抽取了其 64 个月的收益率记录进行审计。
问题: 64 个月样本均值的标准误是多少?
解: $$SE(\bar{x}) = \frac{\sigma}{\sqrt{n}} = \frac{3.5\%}{\sqrt{64}} = \frac{3.5\%}{8} = 0.4375\%$$
含义: 这个基金的实际 64 个月平均收益率如果落在 [1.2% − 1.96×0.4375%, 1.2% + 1.96×0.4375%] = [0.34%, 2.06%] 之外,就可以在 5% 显著性水平下质疑其声称的 μ = 1.2%(与后面假设检验联动)。
6.2 客户满意度调查
场景: 一家连锁餐厅的客户评分总体标准差为 15 分。总部想估计全国平均评分,要求标准误不超过 2 分。
问题: 至少需要抽取多少份问卷?
解: $$SE = \frac{\sigma}{\sqrt{n}} \leq 2$$ $$\frac{15}{\sqrt{n}} \leq 2$$ $$\sqrt{n} \geq 7.5$$ $$n \geq 56.25$$
→ 至少需要 57 份问卷
6.3 股票收益的偏态总体
场景: 创业板股票的日收益率分布高度偏态(正偏,有极端涨停日)。但根据 CLT,如果抽取 n=50 只创业板股票计算平均日收益率……
结论: 即使总体严重偏态,n=50 ≥ 30,根据 CLT,这 50 只股票平均收益率的抽样分布仍近似正态。分析师可以放心使用正态分布的分位数做推断。
七、常见误区辨析
误区 1:CLT 说"样本本身会变正态"
❌ 错误: "抽取一个大样本,样本内的数据分布会接近正态。"
✅ 正确: CLT 说的是样本均值的抽样分布趋向正态,不是说单个样本内部的数据分布变正态。
样本内部的数据分布始终反映总体分布的特征。如果你从偏态总体抽 1000 个数据点,这 1000 个数据点依然是偏态的。但如果你抽 1000 次(每次 n=30),这 1000 个样本均值会呈正态。
误区 2:n ≥ 30 是一条铁律
❌ 错误: "n=29 就是不够,n=30 就完全没问题。"
✅ 正确: n=30 只是一个经验法则。实际取决于总体偏态程度。轻度偏态总体 n=20 已够,极度偏态(如保险索赔金额)可能需要 n=100+。
误区 3:总体必须是正态分布才能用 CLT
❌ 错误: "CLT 只适用于正态分布的总体。"
✅ 正确: CLT 的精髓恰恰是总体不需要是正态分布。这才是它被称为"统计学最重要的定理"之一的原因。
八、CLT 与投资实务的关联
| 应用场景 | CLT 的作用 |
|---|---|
| 基金业绩评估 | 用若干年的月度收益率均值推断真实长期收益率 |
| VaR(风险价值)计算 | 假设组合收益率均值服从正态分布,基于 CLT |
| 抽样审计 | 抽取若干笔交易,用样本均值推断总体均值 |
| 行业比较分析 | 抽取行业代表公司,用样本均值代表行业总体特征 |
| 宏观经济预测 | 用样本数据推断总体参数(通胀率、就业率等) |
九、CLT 在 CFA 知识体系中的位置
描述性统计 → 概率论 → 随机变量分布(正态、t、卡方、F)
↓
抽样与估计 ← 本部分
↓
【中心极限定理】← L126(基础)
↓
标准误(L127)
↓
区间估计(L128-129)
↓
假设检验(后续模块)
⚠️ CLT 是整个推断统计学的基石。如果 CLT 不成立,后面所有的置信区间、假设检验、t 检验、回归推断……全部失效。
十、测试题
题 1
中心极限定理的核心结论是:
A. 任何总体的数据都服从正态分布
B. 当样本量足够大时,样本均值的抽样分布近似正态
C. 当样本量足够大时,样本中每个观测值都近似正态
D. 正态总体的样本均值一定服从正态分布,无论样本量多大
题 2
从均值为 100、标准差为 20 的总体中抽取 n=25 的简单随机样本。样本均值的标准误为:
A. 20
B. 4
C. 0.8
D. 5
题 3
某分析师的总体数据严重右偏(如个人年收入)。他抽取了 n=35 的样本。根据 CLT,以下哪项成立?
A. 样本中的 35 个观测值会近似正态分布
B. 样本均值的抽样分布会近似正态
C. 样本均值等于总体均值
D. 样本的标准差等于 σ/√35
题 4(计算题)
某股票策略的日报酬率标准差为 1.8%。若抽取 81 个交易日的收益率计算平均值,该平均值抽样分布的标准差是多少?
A. 1.8%
B. 0.2%
C. 0.022%
D. 0.18%
题 5(概念题)
以下关于中心极限定理的陈述,哪一项是错误的?
A. CLT 要求样本必须是简单随机样本
B. 总体分布偏态越严重,需要更大的样本量才能保证 CLT 效果
C. CLT 保证了样本均值恰好等于总体均值
D. CLT 是现代统计学推断方法的理论基础
十一、答案与解析
【题 1 答案】B — 当样本量足够大时,样本均值的抽样分布近似正态
| 选项 | 判断 | 理由 |
|---|---|---|
| A | ❌ | 总体本身可以是任何分布,CLT 恰恰不要求总体正态 |
| B | ✅ | CLT 的核心陈述 |
| C | ❌ | 这是最常见的混淆——CLT 说的是样本均值,不是样本内的每个观测值 |
| D | ❌ | 虽然这句话本身对(正态总体下 n 再小也是正态),但它不是 CLT 的定义 |
【题 2 答案】B — 4
$$SE(\bar{x}) = \frac{\sigma}{\sqrt{n}} = \frac{20}{\sqrt{25}} = \frac{20}{5} = 4$$
【题 3 答案】B — 样本均值的抽样分布会近似正态
| 选项 | 判断 | 理由 |
|---|---|---|
| A | ❌ | 这就是最常见的误区——CLT 说的是均值的分布,不是样本内的数据分布 |
| B | ✅ | n=35 ≥ 30,CLT 生效 |
| C | ❌ | 样本均值期望等于总体均值(无偏性),但不等于某个具体样本的 $\bar{x}$ 恰好等于 μ |
| D | ❌ | 样本的标准差是 s(样本标准差),不是 σ/√n。σ/√n 是样本均值的标准误 |
【题 4 答案】B — 0.2%
$$SE = \frac{1.8\%}{\sqrt{81}} = \frac{1.8\%}{9} = 0.2\%$$
【题 5 答案】C — CLT 保证了样本均值恰好等于总体均值
| 选项 | 判断 | 理由 |
|---|---|---|
| A | ✅ | CLT 基于简单随机抽样 |
| B | ✅ | 偏态越严重 → 收敛越慢 → 需要更大的 n |
| C | ❌ | CLT 保证的是期望等于 μ,不是"恰好等于" μ。样本均值是随机变量,肯定有抽样误差 |
| D | ✅ | CLT 是推断统计学的基石 |
十二、本课小结
| 要点 | 一句话总结 |
|---|---|
| CLT 定义 | 独立同分布随机变量的均值的抽样分布,当 n 够大时趋向正态 |
| 三个结论 | ① $E(\bar{x}) = \mu$;② $Var(\bar{x}) = \sigma^2/n$;③ 分布趋近正态 |
| 标准误 | $SE = \sigma / \sqrt{n}$,衡量样本均值的精度 |
| 经验法则 | n ≥ 30 → CLT 生效(CFA 标准答案) |
| 核心误区 | CLT 说的是样本均值的分布,不是样本内个体数据的分布 |
📚 下一课 L127:标准误(Standard Error)——标准误 vs 标准差的区别?t 分布何时登场?抽样分布的完整图谱。
Quantitative Methods — Sampling and Estimation · Lesson 2
1. Starting with a Question
Suppose you want to estimate the average P/E ratio of all stocks listed on the Shenzhen Stock Exchange. There are 2,800+ stocks in the population. You randomly select 100 stocks and compute a sample mean $\bar{x}$ = 15.3.
Question: How far is this 15.3 from the true population mean μ? If you repeated this sampling 1,000 times (each time drawing 100 stocks), what distribution would those $\bar{x}$ values follow?
The answer to this question is the Central Limit Theorem (CLT) .
2. Central Limit Theorem: Core Statement
2.1 Definition
Central Limit Theorem (CLT): For a population with mean μ and variance σ², if we draw simple random samples of size n, then as the sample size n becomes sufficiently large, the sampling distribution of the sample mean $\bar{x}$ approaches a normal distribution with mean μ and variance σ²/n.
In mathematical notation:
$$\bar{x} \sim N\left(\mu, \frac{\sigma^2}{n}\right) \quad \text{(when n is sufficiently large)}$$
2.2 Three Core Conclusions of CLT
| # | Conclusion | Meaning |
|---|---|---|
| 1 | Expectation of sample mean = population mean | $E(\bar{x}) = \mu$ — unbiasedness |
| 2 | Variance of sample mean = σ²/n | $Var(\bar{x}) = \frac{\sigma^2}{n}$ — larger sample, smaller variation |
| 3 | Distribution shape converges to normal | No matter what the population distribution is, as long as n is large enough, the distribution of $\bar{x}$ tends toward normality |
🔥 Conclusion #3 is the essence of CLT: The population can be uniform, exponential, or even highly skewed — as long as n is large enough, the distribution of the sample mean approximates normality. This is what makes CLT so powerful.
3. Intuitive Understanding: Dice Experiment
3.1 Distribution of a Single Die
Rolling one fair die, the distribution of outcome X is discrete uniform (1–6, each 1/6):
P(X)
│
│ █ █ █ █ █ █
│ █ █ █ █ █ █
└──┴──┴──┴──┴──┴──
1 2 3 4 5 6 → Not normal at all!
- μ = 3.5, σ² ≈ 2.92
3.2 Rolling Two Dice, Taking the Mean
Rolling 2 dice and computing the mean $\bar{x}_2$:
P(mean)
│ █
│ █████
│ █████████
│ █████████████
└────────────────
1 2 3 4 5 6 → Starting to look triangular
3.3 Rolling 30 Dice, Taking the Mean
Rolling 30 dice and computing the mean $\bar{x}_{30}$:
P(mean)
│ ╭──╮
│ ╱ ╲
│ ╱ ╲
│ ╱ ╲
└────────────────────
2.5 3.0 3.5 4.0 4.5 → Approximately normal!
💡 The population distribution of a die roll is completely non-normal (discrete uniform), yet when n = 30, the distribution of the sample mean is already very close to normal.
4. Key Conditions of CLT
4.1 How Large is "Sufficiently Large"?
| Population Distribution | Minimum Sample Size Required | Reason |
|---|---|---|
| Population near-normal | n ≥ 10 may suffice | Sampling distribution from a normal population is normal regardless |
| Symmetric but non-normal (e.g., uniform) | n ≥ 15~20 | Symmetry aids rapid convergence |
| Skewed (e.g., income distribution, stock returns) | n ≥ 30 | CFA standard convention |
| Extremely skewed (e.g., insurance claim amounts) | n may need to be larger | More observations needed to counteract extreme skewness |
📌 CFA Level 1 exam standard answer: n ≥ 30 is considered "sufficiently large." This is the widely-accepted Rule of Thumb.
4.2 Prerequisites
| Condition | Explanation |
|---|---|
| Simple random sample | Equal probability of selection for each individual, independent draws |
| Finite population variance σ² | Population variance must exist and be finite |
| n sufficiently large | Typically n ≥ 30 |
5. Mathematical Foundation of CLT
5.1 Sampling Distribution
Definition: The distribution of sample means computed from repeated samples of the same size drawn from the same population. This is the "sampling distribution of $\bar{x}$."
5.2 Three Properties
$$\text{Sampling distribution of the sample mean:}$$
| Parameter | Formula | Description |
|---|---|---|
| Mean | $E(\bar{x}) = \mu$ | $\bar{x}$ is an unbiased estimator of μ |
| Variance | $Var(\bar{x}) = \frac{\sigma^2}{n}$ | Larger sample → smaller variance |
| Standard Deviation (Standard Error) | $SE(\bar{x}) = \frac{\sigma}{\sqrt{n}}$ | Standard error is inversely proportional to √n |
5.3 Meaning of Standard Error
Standard Error (SE) is the standard deviation of the sampling distribution of the sample mean. It measures how much $\bar{x}$ varies around μ.
$$\text{Precision} \propto \frac{1}{\sqrt{n}}$$
| Sample Size n | Standard Error (relative to σ) | Precision Improvement |
|---|---|---|
| 10 | σ/√10 ≈ 0.316σ | — |
| 30 | σ/√30 ≈ 0.183σ | 73% better than n=10 |
| 100 | σ/√100 = 0.1σ | 83% better than n=30 |
| 400 | σ/√400 = 0.05σ | 100% better than n=100 |
🔑 Key insight: To double precision (halve the standard error), you need to quadruple the sample size. The cost of precision increases quadratically.
6. Practical Examples
6.1 Fund Performance Analysis
Scenario: A hedge fund claims its monthly return has a mean of 1.2% with a standard deviation of 3.5%. A regulator randomly selects 64 months of the fund's return records for audit.
Question: What is the standard error of the 64-month sample mean?
Solution: $$SE(\bar{x}) = \frac{\sigma}{\sqrt{n}} = \frac{3.5\%}{\sqrt{64}} = \frac{3.5\%}{8} = 0.4375\%$$
Interpretation: If the actual 64-month average return falls outside [1.2% − 1.96×0.4375%, 1.2% + 1.96×0.4375%] = [0.34%, 2.06%], we can challenge the claimed μ = 1.2% at the 5% significance level (foreshadowing hypothesis testing).
6.2 Customer Satisfaction Survey
Scenario: A restaurant chain's customer satisfaction scores have a population standard deviation of 15 points. Headquarters wants to estimate the national average score with a standard error of no more than 2 points.
Question: What is the minimum sample size required?
Solution: $$SE = \frac{\sigma}{\sqrt{n}} \leq 2$$ $$\frac{15}{\sqrt{n}} \leq 2$$ $$\sqrt{n} \geq 7.5$$ $$n \geq 56.25$$
→ At least 57 surveys required
6.3 Skewed Population of Stock Returns
Scenario: The daily returns of ChiNext (创业板) stocks are highly positively skewed (with extreme limit-up days). However, if we draw n = 50 ChiNext stocks and compute the average daily return...
Conclusion: Even though the population is severely skewed, n = 50 ≥ 30. According to CLT, the sampling distribution of the 50-stock mean return will still be approximately normal. Analysts can confidently use normal distribution critical values for inference.
7. Common Misconceptions
Misconception 1: CLT says "the sample itself becomes normal"
❌ Wrong: "Draw a large sample, and the data within the sample will approximate a normal distribution."
✅ Correct: CLT states that the sampling distribution of the sample mean approaches normality — not that individual observations within a single sample become normal.
Data within a single sample always reflects the shape of the population distribution. If you draw 1,000 observations from a skewed population, those 1,000 data points remain skewed. But if you draw 1,000 separate samples (each with n = 30), those 1,000 sample means will be approximately normal.
Misconception 2: n ≥ 30 is a hard rule
❌ Wrong: "n = 29 is insufficient; n = 30 is perfectly fine."
✅ Correct: n = 30 is merely a rule of thumb. The actual requirement depends on the degree of skewness. A mildly skewed population may be fine with n = 20; an extremely skewed one (e.g., insurance claim amounts) may require n = 100+.
Misconception 3: The population must be normal to apply CLT
❌ Wrong: "CLT only applies to normally distributed populations."
✅ Correct: The entire point of CLT is that the population does NOT need to be normally distributed. This is precisely why it is considered one of the most important theorems in statistics.
8. CLT and Investment Practice
| Application | Role of CLT |
|---|---|
| Fund performance evaluation | Use several years of monthly returns to infer the true long-term return |
| VaR (Value at Risk) calculation | Assume portfolio return mean follows a normal distribution, based on CLT |
| Sampling audit | Draw a number of transactions; use sample mean to infer population mean |
| Industry comparative analysis | Select representative companies; use sample mean to represent industry characteristics |
| Macroeconomic forecasting | Use sample data to infer population parameters (inflation rate, employment rate, etc.) |
9. CLT's Position in the CFA Curriculum
Descriptive Statistics → Probability → Random Variable Distributions (Normal, t, Chi-square, F)
↓
Sampling and Estimation ← This section
↓
【Central Limit Theorem】← L126 (Foundation)
↓
Standard Error (L127)
↓
Interval Estimation (L128–L129)
↓
Hypothesis Testing (Upcoming modules)
⚠️ CLT is the cornerstone of inferential statistics. If CLT did not hold, everything that follows — confidence intervals, hypothesis testing, t-tests, regression inference — would all collapse.
10. Practice Questions
Question 1
The core conclusion of the Central Limit Theorem is:
A. Data from any population follows a normal distribution
B. When sample size is sufficiently large, the sampling distribution of the sample mean is approximately normal
C. When sample size is sufficiently large, each observation in the sample approximates normality
D. Sample means from a normal population always follow a normal distribution, regardless of sample size
Question 2
A simple random sample of n = 25 is drawn from a population with mean 100 and standard deviation 20. The standard error of the sample mean is:
A. 20
B. 4
C. 0.8
D. 5
Question 3
An analyst's population data is severely right-skewed (e.g., annual personal income). He draws a sample of n = 35. According to CLT, which of the following holds?
A. The 35 observations in the sample will be approximately normally distributed
B. The sampling distribution of the sample mean will be approximately normal
C. The sample mean equals the population mean
D. The sample standard deviation equals σ/√35
Question 4 (Calculation)
A stock strategy has a daily return standard deviation of 1.8%. If 81 trading days of returns are sampled to compute the average, what is the standard deviation of the sampling distribution of this average?
A. 1.8%
B. 0.2%
C. 0.022%
D. 0.18%
Question 5 (Conceptual)
Which of the following statements about the Central Limit Theorem is incorrect?
A. CLT requires that the sample be a simple random sample
B. The more skewed the population distribution, the larger the sample size needed for CLT to work
C. CLT guarantees that the sample mean exactly equals the population mean
D. CLT is the theoretical foundation of modern statistical inference methods
11. Answers and Explanations
【Answer 1】B — When sample size is sufficiently large, the sampling distribution of the sample mean is approximately normal
| Option | Verdict | Reasoning |
|---|---|---|
| A | ❌ | The population itself can have any distribution — CLT does not require normality |
| B | ✅ | The core statement of CLT |
| C | ❌ | This is the most common confusion — CLT refers to the sample mean, not individual observations |
| D | ❌ | While this statement itself is true (a normal population yields a normal sampling distribution at any n), it is not the definition of CLT |
【Answer 2】B — 4
$$SE(\bar{x}) = \frac{\sigma}{\sqrt{n}} = \frac{20}{\sqrt{25}} = \frac{20}{5} = 4$$
【Answer 3】B — The sampling distribution of the sample mean will be approximately normal
| Option | Verdict | Reasoning |
|---|---|---|
| A | ❌ | This is the classic misconception — CLT concerns the distribution of the mean, not the data within the sample |
| B | ✅ | n = 35 ≥ 30, CLT applies |
| C | ❌ | The expected value of the sample mean equals the population mean (unbiasedness), but a specific $\bar{x}$ will not exactly equal μ |
| D | ❌ | The sample standard deviation is s, not σ/√n. σ/√n is the standard error of the sample mean |
【Answer 4】B — 0.2%
$$SE = \frac{1.8\%}{\sqrt{81}} = \frac{1.8\%}{9} = 0.2\%$$
【Answer 5】C — CLT guarantees that the sample mean exactly equals the population mean
| Option | Verdict | Reasoning |
|---|---|---|
| A | ✅ | CLT is based on simple random sampling |
| B | ✅ | Greater skewness → slower convergence → larger n required |
| C | ❌ | CLT guarantees that the expectation equals μ, not that any particular sample mean "exactly equals" μ. The sample mean is a random variable and will always have sampling error |
| D | ✅ | CLT is the cornerstone of inferential statistics |
12. Summary
| Key Point | One-Liner |
|---|---|
| CLT Definition | The sampling distribution of the mean of i.i.d. random variables approaches normality as n increases |
| Three Conclusions | ① $E(\bar{x}) = \mu$; ② $Var(\bar{x}) = \sigma^2/n$; ③ Distribution converges to normal |
| Standard Error | $SE = \sigma / \sqrt{n}$, measures the precision of the sample mean |
| Rule of Thumb | n ≥ 30 → CLT applies (CFA standard answer) |
| Core Misconception | CLT is about the distribution of sample means, not the distribution of individual data within a sample |
📚 Next Lesson L127: Standard Error — Standard Error vs Standard Deviation? When does the t-distribution enter the picture? Complete mapping of sampling distributions.