E(X) generalises the plain numerical average, weighting each value by its true probability rather than by n1; it need not be a value X can actually take, and is best read as the long-run average over many repetitions. Theorem 11.3 extends this to any function g(X): E(g(X))=∑xg(x)f(x) or ∫g(x)f(x)dx; taking g(X)=Xk gives the k-th momentE(Xk).
Variance (Definition 11.9): V(X)=E((X−E(X))2), with the far more usable computing form
V(X)=E(X2)−(E(X))2.
Standard deviation is σ=V(X); both are always ≥0. A smaller σ2 means values cluster tightly around the mean; a larger σ2 means they scatter more widely — even distributions sharing the same mean can differ sharply here.
Three linearity laws (for constants a,b): E(aX+b)=aE(X)+b (so E(aX)=aE(X) and E(b)=b); V(X)=E(X2)−(E(X))2 (restated); and V(aX+b)=a2V(X) (so V(aX)=a2V(X) and V(b)=0). These make quick work of a shifted/scaled random variable — e.g. a net "winning amount" that is a linear function of a raw count — without recomputing the distribution from scratch.
Worked technique. For a discrete X: tabulate x, f(x), xf(x), x2f(x); sum the last two columns to get E(X) and E(X2) directly, then apply V(X)=E(X2)−(E(X))2. For a continuous X: compute E(X)=∫xf(x)dx and E(X2)=∫x2f(x)dx over the support, then the same variance formula.
E(X+3)=E(X)+3=10⇒μ=7; expand E((X+3)2)=E(X2)+6E(X)+9=116 to get E(X2)=65, then σ2=E(X2)−μ2.
✓Final answer
μ=7,σ2=16.
Use linearity of expectation on E(X+3) to get μ directly, then expand E((X+3)2) into moments of X to get E(X2), and finish with σ2=E(X2)−μ2.
Step 1. Find μ=E(X).E(X+3)=E(X)+3=10⇒E(X)=7, so μ=7.
Step 2. Expand E((X+3)2).(X+3)2=X2+6X+9, so E((X+3)2)=E(X2)+6E(X)+9.
Step 3. Substitute the given value and solve for E(X2).E(X2)+6(7)+9=116⇒E(X2)+42+9=116⇒E(X2)=116−51=65.
Step 4. Compute the variance.σ2=E(X2)−μ2=65−72=65−49=16.
✓Final answer
μ=7 and σ2=16.
Linearity of E on a shifted variable, then σ2=E(X2)−μ2
Expanding (X+3)2 incorrectly as X2+9 (dropping the cross term 6X)
Using E(X+3) itself as μ instead of subtracting the 3