The study of many phenomena often involves collecting numerical values representing objective measurements of a series of features that characterize the phenomena in question. Here we consider the simplified case in which a single feature is measured across a certain number, n, of statistical units (objects from which a measurement can be obtained). This produces the numerical sequence x1, x2… *xn, corresponding to the n measurements. Examples of this situation are common in applications. We then seek to summarize this collection of numbers "as effectively as possible" by introducing both a central value and a measure of dispersion. Since the work of Carl Friedrich Gauss (1777–1855), the usual choices have been the arithmetic mean as the central value and the variance* as the measure of dispersion. The first of these parameters minimizes the sum of the squared deviations between the measurements and any constant value; the second distributes this minimum total deviation across the number of observations.
Mean and variance -------------------
Let m denote the mean and s2 the variance. By definition, we can write:
m=1ni=1nxiands2=1ni=1n(sim)2.m=\frac{1}{n}\sum_{i=1}^{n}x_i\quad\text{and}\quad s^2=\frac{1}{n}\sum_{i=1}^{n}(s_i-m)^2 .
The variance does not have the same unit of measurement as the variable under study (if m is measured in meters, s2 is measured in square meters; if m is expressed in euros, s2 is expressed in "square euros," etc.). To recover the original unit, we introduce the standard deviation s, defined as the (positive) square root of the variance.