定量方法(Quantitative Methods)— 简单线性回归 · 第 2 课
一、本课定位
| 课次 | 主题 | 核心能力 |
|---|---|---|
| L137 | 简单线性回归模型 | 理解 Y = b0 + b1X + ε 的结构与含义 |
| L138 | 最小二乘法(OLS) | 学会计算 b0、b1,让误差平方和最小 |
| L139 | R² 与 F 检验 | 模型整体解释力 |
| L140 | 回归假设与诊断 | 检验假设是否成立 |
| L141 | 定量模块终测 | 15 题综合测试 |
🎯 L138 是「从读懂模型」到「亲手算出模型」的关键一课——上一课你会读表,这一课你要会算数。
二、我们要解决什么问题?
上一课我们知道: 回归线是 Y = b0 + b1X,但我们不知道 b0、b1 具体是多少。
L138 的核心任务: 给定一组数据 (X1,Y1), (X2,Y2), ..., (Xn,Yn),找到「最好的」一条直线。
Y ↑
| · · ← 观测到的点
| · · ·
| · ╱ ← 我们要找的那条直线
| ╱ ╱
| ╱ ╱
+──────────────→ X
什么是「最好」? —— 让所有点到直线的「垂直距离」整体最小。
三、普通最小二乘法(Ordinary Least Squares, OLS)
3.1 核心思想
OLS 找那条让「残差平方和」最小的直线。
- 残差(Residual) = 实际值 − 拟合值 = Yi − Ŷi = Yi − (b0 + b1Xi)
- 残差平方和(SSE / SSR) = Σ(Yi − Ŷi)² = Σ ε̂i²
OLS 的目标函数:
$$\min_{b_0, b_1} \sum_{i=1}^{n} \left(Y_i - b_0 - b_1 X_i\right)^2$$
3.2 为什么要「平方」?
| 做法 | 问题 |
|---|---|
| 直接加总残差 Σεi | 正负抵消,可能等于 0,没有信息量 |
| 加总绝对值 Σ|εi| | 数学上不好求导(有尖点) |
| 加总平方 Σεi² | ✅ 全部变正 + 惩罚大误差 + 可求导 |
平方的意义:① 消除正负号;② 大误差受到更大惩罚;③ 数学上可微,能求出唯一解。
3.3 「普通」二字从哪来?
| 方法 | 特点 |
|---|---|
| 普通最小二乘 OLS | 每个点权重相同(最常用) |
| 加权最小二乘 WLS | 给精度高的点更大权重 |
| 广义最小二乘 GLS | 处理异方差/自相关 |
CFA 一级只考 OLS。记住「普通」= 所有观测一视同仁。
四、OLS 的求解公式(🔑 必背)
通过让 SSE 对 b0、b1 分别求导并令其等于 0,得到:
4.1 斜率 b₁
$$b_1 = \frac{\sum(X_i - \bar{X})(Y_i - \bar{Y})}{\sum(X_i - \bar{X})^2} = \frac{\text{Cov}(X,Y)}{\text{Var}(X)}$$
4.2 截距 b₀
$$b_0 = \bar{Y} - b_1 \bar{X}$$
符号说明:
| 符号 | 含义 |
|---|---|
| X̄ | X 的样本均值 |
| Ȳ | Y 的样本均值 |
| Cov(X,Y) | X 与 Y 的样本协方差 |
| Var(X) | X 的样本方差 |
4.3 两个重要性质
- 回归线一定经过均值点 (X̄, Ȳ) —— 这是 b₀ 公式的直接推论。
- 残差的均值为 0:Σε̂i = 0(OLS 必然满足,正负残差恰好平衡)。
📌 记忆口诀:b₁ = 协方差 ÷ 方差;b₀ = Ȳ − b₁X̄。
五、完整计算示例(手把手)
数据: 用 5 个月的广告费(X,万元)预测销售额(Y,万元)
| 月份 | X(广告费) | Y(销售额) |
|---|---|---|
| 1 | 1 | 2 |
| 2 | 2 | 4 |
| 3 | 3 | 5 |
| 4 | 4 | 4 |
| 5 | 5 | 10 |
5.1 先算均值
- X̄ = (1+2+3+4+5)/5 = 15/5 = 3
- Ȳ = (2+4+5+4+10)/5 = 25/5 = 5
5.2 列表计算离差
| i | Xi | Yi | Xi−X̄ | Yi−Ȳ | (Xi−X̄)(Yi−Ȳ) | (Xi−X̄)² |
|---|---|---|---|---|---|---|
| 1 | 1 | 2 | −2 | −3 | 6 | 4 |
| 2 | 2 | 4 | −1 | −1 | 1 | 1 |
| 3 | 3 | 5 | 0 | 0 | 0 | 0 |
| 4 | 4 | 4 | 1 | −1 | −1 | 1 |
| 5 | 5 | 10 | 2 | 5 | 10 | 4 |
| Σ | 16 | 10 |
5.3 代入公式
$$b_1 = \frac{16}{10} = 1.6$$
$$b_0 = \bar{Y} - b_1\bar{X} = 5 - 1.6 \times 3 = 5 - 4.8 = 0.2$$
5.4 结果
$$\hat{Y} = 0.2 + 1.6X$$
解释:每多投 1 万元广告费,销售额预期增加 1.6 万元;不投广告时,基础销售额约 0.2 万元。
六、残差与拟合值(看懂模型的「误差」)
用上面的模型 Ŷ = 0.2 + 1.6X,回算每个点:
| i | Xi | 实际 Yi | 拟合值 Ŷi = 0.2+1.6Xi | 残差 ε̂i = Yi−Ŷi | 残差² |
|---|---|---|---|---|---|
| 1 | 1 | 2 | 1.8 | 0.2 | 0.04 |
| 2 | 2 | 4 | 3.4 | 0.6 | 0.36 |
| 3 | 3 | 5 | 5.0 | 0.0 | 0.00 |
| 4 | 4 | 4 | 6.6 | −2.6 | 6.76 |
| 5 | 5 | 10 | 8.2 | 1.8 | 3.24 |
| Σ | 0 | 10.4 |
验证:① 残差之和 = 0 ✅;② SSE = 10.4,这是所有可能直线中「最小的平方和」——OLS 保证这一点。
七、估计标准误(Standard Error of Estimate, SEE)
衡量模型预测误差的典型大小:
$$SEE = \sqrt{\frac{SSE}{n - k - 1}} = \sqrt{\frac{\sum \hat{\varepsilon}_i^2}{n - 2}}$$
| 符号 | 含义 |
|---|---|
| n | 样本量 |
| k | 自变量个数(简单回归 k=1) |
| n − k − 1 | 自由度(简单回归 = n − 2) |
简单线性回归中自由度 = n − 2,因为我们已经用数据估计了 b0、b1 两个参数。
SEE 越小 → 点离回归线越近 → 模型拟合越好。
八、练习题(10 题)
基础概念(Q1–Q5)
Q1. OLS 最小化的量是: A. 残差之和 Σεi B. 残差平方和 Σεi² C. 因变量之和 ΣYi D. 自变量之和 ΣXi
Q2. 残差(Residual)的定义是: A. 拟合值减实际值 B. 实际值减拟合值 C. 实际值减均值 D. 拟合值减均值
Q3. OLS 斜率 b₁ 的计算公式是: A. Var(X)/Cov(X,Y) B. Cov(X,Y)/Var(X) C. Cov(X,Y)/Var(Y) D. Var(Y)/Cov(X,Y)
Q4. 关于 OLS 回归线,一定成立的是: A. 经过原点 (0,0) B. 经过均值点 (X̄, Ȳ) C. 残差平方和等于 0 D. 斜率一定为正
Q5. OLS 中「普通」的含义是: A. 模型只有一个自变量 B. 所有观测点权重相同 C. 不需要假设条件 D. 结果一定显著
计算应用(Q6–Q8)
Q6. 已知 X̄=10, Ȳ=20, b₁=0.5,则截距 b₀ 等于: A. 10 B. 15 C. 20 D. 25
Q7. 某样本 Σ(Xi−X̄)(Yi−Ȳ)=40,Σ(Xi−X̄)²=20,则 b₁ 等于: A. 0.5 B. 2.0 C. 20 D. 40
Q8. 一个回归的 SSE=30,样本量 n=8,简单线性回归的 SEE 等于: A. √(30/6) B. √(30/7) C. √(30/8) D. 30/6
综合推理(Q9–Q10)
Q9. 回归模型 Ŷ = 3 + 2X,当 X=4 时实际 Y=12,则残差为: A. −1 B. 1 C. 2 D. 11
Q10. 下列关于 OLS 残差的性质,正确的是: A. 残差之和恒为正 B. 残差平方和最小,但残差之和不一定为 0 C. 残差之和恒为 0,且残差平方和最小 D. 残差之和恒等于 n
九、答案与详解
| 题号 | 答案 | 详解 |
|---|---|---|
| Q1 | B | OLS = Ordinary Least Squares,最小化的就是「平方和」Σεi²。A 会正负抵消。 |
| Q2 | B | 残差 = 实际 Yi − 拟合 Ŷi。A 方向反了。C/D 是离差,不是残差。 |
| Q3 | B | b₁ = Cov(X,Y)/Var(X) = Σ(Xi−X̄)(Yi−Ȳ)/Σ(Xi−X̄)²。 |
| Q4 | B | b₀ = Ȳ − b₁X̄ 直接保证直线过 (X̄,Ȳ)。A 错:只有 b₀=0 时才过原点。C 错:残差平方和一般不为 0。 |
| Q5 | B | 「普通」= 每个观测同等权重,区别于 WLS/GLS。 |
| Q6 | B | b₀ = Ȳ − b₁X̄ = 20 − 0.5×10 = 20 − 5 = 15。 |
| Q7 | B | b₁ = 40/20 = 2.0。 |
| Q8 | A | SEE = √(SSE/(n−2)) = √(30/6)。简单回归自由度 n−2=6。 |
| Q9 | B | Ŷ = 3 + 2×4 = 11;残差 = 12 − 11 = 1。 |
| Q10 | C | OLS 的两个经典性质:残差之和 = 0(正负平衡),且残差平方和最小(最小二乘的定义)。 |
十、CFA 一级核心考点
| 考点 | 记忆要点 |
|---|---|
| OLS 目标 | 最小化残差平方和 Σ(Yi−Ŷi)² |
| 斜率公式 | b₁ = Cov(X,Y)/Var(X) |
| 截距公式 | b₀ = Ȳ − b₁X̄ |
| 过均值点 | 回归线必过 (X̄, Ȳ) |
| 残差性质 | Σε̂i = 0 且平方和最小 |
| 自由度 | 简单回归 = n − 2(估计了 b0、b1 两个参数) |
| SEE | √(SSE/(n−2)),越小拟合越好 |
📌 下一课 L139:R² 与 F 检验——这条线到底解释了 Y 多少波动?怎么判断模型整体有没有用?
Quantitative Methods — Simple Linear Regression · Lesson 2
I. Where This Lesson Fits
| Lesson | Topic | Core Skill |
|---|---|---|
| L137 | Simple Linear Regression Model | Understand the structure and meaning of Y = b0 + b1X + ε |
| L138 | Ordinary Least Squares (OLS) | Learn to compute b0 and b1 by minimizing the sum of squared errors |
| L139 | R² and the F-Test | Overall explanatory power of the model |
| L140 | Regression Assumptions & Diagnostics | Test whether assumptions hold |
| L141 | Quant Methods Final Quiz | 15-question comprehensive test |
🎯 L138 is the bridge from "reading the model" to "computing the model yourself" — last lesson you read a table, this lesson you do the math.
II. What Problem Are We Solving?
From last lesson: a regression line is Y = b0 + b1X, but we don't yet know the actual values of b0 and b1.
Core task of L138: given a set of data points (X1,Y1), (X2,Y2), ..., (Xn,Yn), find the "best" straight line.
Y ↑
| · · ← observed points
| · · ·
| · ╱ ← the line we're looking for
| ╱ ╱
| ╱ ╱
+──────────────→ X
What does "best" mean? — Make the overall vertical distances from all points to the line as small as possible.
III. Ordinary Least Squares (OLS)
3.1 The Core Idea
OLS finds the line that minimizes the sum of squared residuals.
- Residual = Actual value − Fitted value = Yi − Ŷi = Yi − (b0 + b1Xi)
- Sum of Squared Errors (SSE / SSR) = Σ(Yi − Ŷi)² = Σ ε̂i²
OLS objective function:
$$\min_{b_0, b_1} \sum_{i=1}^{n} \left(Y_i - b_0 - b_1 X_i\right)^2$$
3.2 Why Square?
| Approach | Problem |
|---|---|
| Sum residuals Σεi | Positives and negatives cancel out; may equal 0 with no information |
| Sum absolute values Σ|εi| | Hard to differentiate (kink at zero) |
| Sum squares Σεi² | ✅ All positive + penalizes large errors + differentiable |
Why squaring works: ① removes the sign; ② large errors get penalized more; ③ mathematically differentiable so a unique solution exists.
3.3 Where Does "Ordinary" Come From?
| Method | Feature |
|---|---|
| Ordinary Least Squares (OLS) | Every point gets equal weight (most common) |
| Weighted Least Squares (WLS) | More weight to more precise points |
| Generalized Least Squares (GLS) | Handles heteroskedasticity / autocorrelation |
CFA Level 1 only tests OLS. Remember: "ordinary" = all observations treated equally.
IV. OLS Formulas (🔑 Must Memorize)
Setting the derivatives of SSE with respect to b0 and b1 to zero yields:
4.1 Slope b₁
$$b_1 = \frac{\sum(X_i - \bar{X})(Y_i - \bar{Y})}{\sum(X_i - \bar{X})^2} = \frac{\text{Cov}(X,Y)}{\text{Var}(X)}$$
4.2 Intercept b₀
$$b_0 = \bar{Y} - b_1 \bar{X}$$
Notation:
| Symbol | Meaning |
|---|---|
| X̄ | Sample mean of X |
| Ȳ | Sample mean of Y |
| Cov(X,Y) | Sample covariance of X and Y |
| Var(X) | Sample variance of X |
4.3 Two Key Properties
- The regression line always passes through the mean point (X̄, Ȳ) — a direct consequence of the b₀ formula.
- Residuals sum to zero: Σε̂i = 0 (OLS guarantees positive and negative residuals exactly balance).
📌 Memory aid: b₁ = covariance ÷ variance; b₀ = Ȳ − b₁X̄.
V. Complete Worked Example (Step by Step)
Data: Use 5 months of advertising spend (X, in $10k) to predict sales (Y, in $10k)
| Month | X (Ad spend) | Y (Sales) |
|---|---|---|
| 1 | 1 | 2 |
| 2 | 2 | 4 |
| 3 | 3 | 5 |
| 4 | 4 | 4 |
| 5 | 5 | 10 |
5.1 Compute the Means
- X̄ = (1+2+3+4+5)/5 = 15/5 = 3
- Ȳ = (2+4+5+4+10)/5 = 25/5 = 5
5.2 Compute Deviations
| i | Xi | Yi | Xi−X̄ | Yi−Ȳ | (Xi−X̄)(Yi−Ȳ) | (Xi−X̄)² |
|---|---|---|---|---|---|---|
| 1 | 1 | 2 | −2 | −3 | 6 | 4 |
| 2 | 2 | 4 | −1 | −1 | 1 | 1 |
| 3 | 3 | 5 | 0 | 0 | 0 | 0 |
| 4 | 4 | 4 | 1 | −1 | −1 | 1 |
| 5 | 5 | 10 | 2 | 5 | 10 | 4 |
| Σ | 16 | 10 |
5.3 Plug Into the Formulas
$$b_1 = \frac{16}{10} = 1.6$$
$$b_0 = \bar{Y} - b_1\bar{X} = 5 - 1.6 \times 3 = 5 - 4.8 = 0.2$$
5.4 Result
$$\hat{Y} = 0.2 + 1.6X$$
Interpretation: each additional $10k of advertising is expected to raise sales by 1.6 ($10k); with zero advertising, baseline sales ≈ 0.2 ($10k).
VI. Residuals and Fitted Values (Understanding the Model's "Error")
Using Ŷ = 0.2 + 1.6X, compute each fitted value:
| i | Xi | Actual Yi | Fitted Ŷi = 0.2+1.6Xi | Residual ε̂i = Yi−Ŷi | Residual² |
|---|---|---|---|---|---|
| 1 | 1 | 2 | 1.8 | 0.2 | 0.04 |
| 2 | 2 | 4 | 3.4 | 0.6 | 0.36 |
| 3 | 3 | 5 | 5.0 | 0.0 | 0.00 |
| 4 | 4 | 4 | 6.6 | −2.6 | 6.76 |
| 5 | 5 | 10 | 8.2 | 1.8 | 3.24 |
| Σ | 0 | 10.4 |
Verification: ① residuals sum to 0 ✅; ② SSE = 10.4, which is the smallest possible squared sum among all candidate lines — guaranteed by OLS.
VII. Standard Error of Estimate (SEE)
Measures the typical size of the model's prediction error:
$$SEE = \sqrt{\frac{SSE}{n - k - 1}} = \sqrt{\frac{\sum \hat{\varepsilon}_i^2}{n - 2}}$$
| Symbol | Meaning |
|---|---|
| n | Sample size |
| k | Number of independent variables (k=1 in simple regression) |
| n − k − 1 | Degrees of freedom (n − 2 in simple regression) |
In simple linear regression the degrees of freedom = n − 2, because we already used the data to estimate two parameters (b0, b1).
Smaller SEE → points closer to the line → better model fit.
VIII. Practice Questions (10 Questions)
Basic Concepts (Q1–Q5)
Q1. OLS minimizes: A. Sum of residuals Σεi B. Sum of squared residuals Σεi² C. Sum of dependent variable ΣYi D. Sum of independent variable ΣXi
Q2. A residual is defined as: A. Fitted value minus actual value B. Actual value minus fitted value C. Actual value minus the mean D. Fitted value minus the mean
Q3. The OLS slope b₁ is computed as: A. Var(X)/Cov(X,Y) B. Cov(X,Y)/Var(X) C. Cov(X,Y)/Var(Y) D. Var(Y)/Cov(X,Y)
Q4. Which is always true for an OLS regression line? A. It passes through the origin (0,0) B. It passes through the mean point (X̄, Ȳ) C. The sum of squared residuals equals 0 D. The slope is always positive
Q5. "Ordinary" in OLS means: A. The model has only one independent variable B. All observations are weighted equally C. No assumptions are required D. The result is always significant
Computation & Application (Q6–Q8)
Q6. Given X̄=10, Ȳ=20, b₁=0.5, the intercept b₀ equals: A. 10 B. 15 C. 20 D. 25
Q7. A sample has Σ(Xi−X̄)(Yi−Ȳ)=40 and Σ(Xi−X̄)²=20. Then b₁ equals: A. 0.5 B. 2.0 C. 20 D. 40
Q8. A regression has SSE=30 with sample size n=8. The SEE for simple linear regression is: A. √(30/6) B. √(30/7) C. √(30/8) D. 30/6
Integrated Reasoning (Q9–Q10)
Q9. For the model Ŷ = 3 + 2X, when X=4 the actual Y=12. The residual is: A. −1 B. 1 C. 2 D. 11
Q10. Which statement about OLS residuals is correct? A. Residuals always sum to a positive value B. The sum of squared residuals is minimized, but residuals may not sum to 0 C. Residuals always sum to 0, and the sum of squared residuals is minimized D. Residuals always sum to n
IX. Answers & Explanations
| Q | Answer | Explanation |
|---|---|---|
| Q1 | B | OLS = Ordinary Least Squares; it minimizes the sum of "squares" Σεi². A cancels out. |
| Q2 | B | Residual = actual Yi − fitted Ŷi. A is reversed. C/D are deviations, not residuals. |
| Q3 | B | b₁ = Cov(X,Y)/Var(X) = Σ(Xi−X̄)(Yi−Ȳ)/Σ(Xi−X̄)². |
| Q4 | B | b₀ = Ȳ − b₁X̄ directly guarantees the line passes through (X̄,Ȳ). A is wrong: it passes through the origin only if b₀=0. C is wrong: SSE is generally not 0. |
| Q5 | B | "Ordinary" = every observation weighted equally, as opposed to WLS/GLS. |
| Q6 | B | b₀ = Ȳ − b₁X̄ = 20 − 0.5×10 = 20 − 5 = 15. |
| Q7 | B | b₁ = 40/20 = 2.0. |
| Q8 | A | SEE = √(SSE/(n−2)) = √(30/6). Simple regression df = n−2 = 6. |
| Q9 | B | Ŷ = 3 + 2×4 = 11; residual = 12 − 11 = 1. |
| Q10 | C | Two classic OLS properties: residuals sum to 0 (balanced), and the sum of squared residuals is minimized (definition of least squares). |
X. CFA Level 1 Key Takeaways
| Concept | Memory Point |
|---|---|
| OLS objective | Minimize the sum of squared residuals Σ(Yi−Ŷi)² |
| Slope formula | b₁ = Cov(X,Y)/Var(X) |
| Intercept formula | b₀ = Ȳ − b₁X̄ |
| Passes through mean | The regression line always passes through (X̄, Ȳ) |
| Residual properties | Σε̂i = 0 and squared sum is minimized |
| Degrees of freedom | Simple regression = n − 2 (two parameters b0, b1 estimated) |
| SEE | √(SSE/(n−2)); smaller = better fit |
📌 Next lesson L139: R² and the F-Test — how much of Y's variation does this line actually explain? And how do we test whether the model is useful at all?