Mathematics
Descriptive statistics calculator
Paste or type your data: mean, median, mode, standard deviation and quartiles are calculated as you write.
How to read these measures
Descriptive statistics summarise a set of data in a handful of measures, and they fall into two families. Measures of location say where the centre is: mean, median and mode. Measures of dispersion say how far the values stray from that centre: range, variance, standard deviation and interquartile range. You need both, because two sets with the same mean can have completely different shapes.
Mean and median answer the same question in different ways. The mean adds up every value and divides by how many there are, so it is pulled by extreme values: a single very high salary shifts the average for an office. The median is the value in the middle once the data are sorted, and one outlier does not move it. When mean and median differ a lot, the distribution is skewed and it is the median that better describes the typical case.
The standard deviation is the square root of the variance, and it has the advantage of being expressed in the same units as the data. There are two versions: the population one divides by n and is used when the data are the whole group of interest; the sample one divides by n − 1 (Bessel's correction) and is used when the data are a sample drawn from a wider population. The second is slightly larger, because it corrects the tendency of a sample to understate the real variability.
A worked example with the data 12, 15, 15, 18, 21, 24, 24, 24, 30: the sum is 183 across 9 values, so the mean is 20.33. Sorted, the middle value is the fifth, 21: that is the median. 24 appears three times and is the mode. The population standard deviation is 5.52, the sample one 5.85. The first quartile is 15 and the third is 24, so half the data sit between those two values.
Common mistakes
- Using the sample standard deviation for data that represent the whole population, or the other way round: the difference is negligible on large sets but noticeable below about thirty observations.
- Describing a strongly skewed distribution, such as incomes, by the mean alone: the median tells you far more about the typical case.
- Confusing no mode with multiple modes: if every value appears the same number of times there is no mode, while two values tied at the highest frequency give a bimodal distribution.
- Comparing standard deviations of quantities with different units or orders of magnitude: that is what the coefficient of variation is for, since it scales dispersion by the mean.
- Forgetting that quartiles can be computed under different conventions: this page uses linear interpolation, the same as spreadsheets, and other textbooks may return slightly different values.
Frequently asked questions
What is the difference between mean, median and mode?
The mean is the sum of the values divided by how many there are. The median is the middle value of the sorted data, or the average of the two middle ones if there is an even count. The mode is the value that occurs most often. On a symmetric distribution they nearly coincide; on a skewed one they separate.
Should I use the population or the sample standard deviation?
If the data are the entire group you care about — the marks of a whole class — use the population one, which divides by n. If they are a sample from which you want to draw conclusions about a wider group, use the sample one, which divides by n − 1.
What does the coefficient of variation tell you?
It is the ratio of the standard deviation to the mean, expressed as a percentage. It lets you compare the dispersion of different quantities: a variability of 5 around a mean of 10 matters far more than the same variability around a mean of 1,000.
How are quartiles calculated?
Q1 leaves a quarter of the sorted data below it, Q3 three quarters. This calculator obtains them by linear interpolation between the values adjacent to the theoretical position, the same convention spreadsheets use. The difference Q3 − Q1 is the interquartile range.
What format should the data be in?
Whichever you prefer: separated by spaces, commas, semicolons or line breaks. You can paste a column copied straight from Excel or a CSV. A comma is treated as a separator between values, so use a full stop for decimals.
How this calculation works
Mean: x̄ = Σxᵢ / n. Median: the value at position (n + 1) / 2 in the sorted data, interpolated when that position is not a whole number. Population variance: σ² = Σ(xᵢ − x̄)² / n; sample variance: s² = Σ(xᵢ − x̄)² / (n − 1). The standard deviation is the square root of the respective variance. Quartiles are taken at position p × (n − 1) with linear interpolation between adjacent values. Coefficient of variation = σ / |x̄| × 100.
Related calculators
Error propagation
The uncertainty of a quantity computed from any formula, in quadrature and as a maximum error, and the mean of repeated readings with its error.
Linear regression
The least-squares line with the uncertainties of slope and intercept, weighted when you have σ, correlation coefficient and χ², with the plot.
Grade average
Simple and weighted averages, and the mark needed for a target.
Probability
Classical and conditional probability, Bayes and the binomial distribution.