The study of many phenomena often involves collecting numerical values representing objective measurements of a series of features that characterize the phenomena in question. Here we consider the simplified case in which a single feature is measured across a certain number, n, of statistical units (objects from which a measurement can be obtained). This produces the numerical sequence x1, x2… *xn, corresponding to the n measurements. Examples of this situation are common in applications. We then seek to summarize this collection of numbers "as effectively as possible" by introducing both a central value and a measure of dispersion. Since the work of Carl Friedrich Gauss (1777–1855), the usual choices have been the arithmetic mean as the central value and the variance* as the measure of dispersion. The first of these parameters minimizes the sum of the squared deviations between the measurements and any constant value; the second distributes this minimum total deviation across the number of observations.
Mean and variance
-------------------
Let m denote the mean and s2 the variance. By definition, we can write:
m=n1∑i=1nxiands2=n1∑i=1n(si−m)2.
The variance does not have the same unit of measurement as the variable under study (if m is measured in meters, s2 is measured in square meters; if m is expressed in euros, s2 is expressed in "square euros," etc.). To recover the original unit, we introduce the standard deviation s, defined as the (positive) square root of the variance.