E(X) generalises the plain numerical average, weighting each value by its true probability rather than by n1; it need not be a value X can actually take, and is best read as the long-run average over many repetitions. Theorem 11.3 extends this to any function g(X): E(g(X))=∑xg(x)f(x) or ∫g(x)f(x)dx; taking g(X)=Xk gives the k-th momentE(Xk).
Variance (Definition 11.9): V(X)=E((X−E(X))2), with the far more usable computing form
V(X)=E(X2)−(E(X))2.
Standard deviation is σ=V(X); both are always ≥0. A smaller σ2 means values cluster tightly around the mean; a larger σ2 means they scatter more widely — even distributions sharing the same mean can differ sharply here. …