Math
Sample Size
Sample Size summarizes a batch of numbers.
What this does
Sample Size summarizes a batch of numbers.
Sample Size summarizes a batch of numbers so you can see the shape of data at a glance. Raw lists are hard to read: fifty test scores tell you little until you know the center, the spread, and whether anything looks unusual. Descriptive statistics turn the list into a handful of honest summary numbers.
The formula, explained plainly
Center has three standard measures. The mean is the balance point: add everything and divide by the count. The median is the middle value after sorting, immune to extremes. The mode is the most frequent value, the only center that works for categories like favorite colors. Each answers a slightly different question, so the right choice depends on the data's shape.
Spread tells you how far values typically stray from the center. The range is the gap between max and min. Variance averages the squared distances from the mean, and the standard deviation is its square root, back in the original units. A small standard deviation means values cluster tightly; a large one means they scatter widely.
Probability tools in this group answer 'how likely' questions with counting, not gut feeling. A fair coin lands heads with probability 1/2 because one of two equally likely outcomes is heads. Combinations and permutations count the possible outcomes first, then divide the favorable ones by the total, which is the entire logic of odds in one sentence.
Distributions describe how values pile up. The normal (bell) curve is the famous one: symmetric, with about 68% of values within one standard deviation of the mean. Percentiles and z-scores then say exactly where one value sits relative to the pack. Sample Size computes these summaries exactly from the numbers you enter.
How to use it
- Enter margin of error (%).
- Enter std deviation.
- Read the instant result and the breakdown below it.
- Adjust any input to compare scenarios.
Worked example
With margin of error (%) = 5, std deviation = 0.5, the result is 385 samples. The default list is small enough to check the arithmetic manually.
Common mistakes
- Quoting percentages from tiny samples as if they were stable: 2 out of 3 is 67%, but one more observation swings it wildly.
- Treating probability as a promise: a 90% chance still fails one time in ten, and streaks are normal in random sequences.
- Confusing percentile with percent: the 90th percentile is a rank (you beat 90% of the group), not a score of 90%.
- Assuming data is normally distributed just because the bell curve is famous; incomes, wait times, and city sizes are heavily skewed.
- Cherry-picking the date range or subgroup that supports a conclusion and ignoring the rest of the data.
- Mixing units inside one dataset, like combining centimeters and inches before computing the mean.
- Reporting an average without any measure of spread, which hides whether the data clusters tightly or scatters everywhere.
- Trusting the mean on skewed data: a few billionaires make the 'average' income meaningless, which is exactly when the median earns its keep.
Limitations
- Models like the normal distribution are approximations; real data has skew, outliers, and lumps the model smooths over.
- Probability here is mathematical (counting equally likely outcomes), not a prediction about any single real event.
- Descriptive statistics summarize data; they never explain why the numbers look that way.
- Garbage in, garbage out: a typo, a wrong unit, or a biased sample corrupts every summary number equally.
- Most inference tools assume roughly random, independent samples; convenience samples and self-selected polls violate that.
Expected accuracy
All summaries are computed exactly from the values you enter using standard statistical definitions, in double-precision arithmetic. Population vs sample variants follow the standard N vs N-1 convention and are labeled where the tool offers both. Results are only as good as the input data: the math cannot detect typos or biased samples.
Privacy
Everything you type stays on your device. The calculation runs in your browser with JavaScript; no input is sent to a server, stored in an account, or shared with anyone.
Sources and standards
- Standard statistical definitions (mean, median, mode, variance, standard deviation, percentiles, z-scores) as used in introductory statistics and the NIST/SEMATECH e-Handbook of Statistical Methods; probability follows classical counting definitions.
Bottom line
Sample Size reduces any list of numbers to its honest essentials: where the center sits, how far values spread, and how likely outcomes are. Prefer the median when outliers lurk, always pair an average with a spread, and remember that summaries describe data without explaining it. Enter values separated by commas and let the math do the reading.
Key insight
The average hides the shape. Before trusting any average, ask whether the data is skewed, because one billionaire in a room makes the average net worth meaningless for everyone else in it.
Frequently asked questions
What is a weighted mean?
A mean where some values count more than others. A course grade might weight the final exam at 40% and homework at 60%. Multiply each value by its weight, add them up, and divide by the total weight. Use it whenever the inputs genuinely deserve unequal influence.
Why do random streaks happen?
Because randomness has no memory. Five heads in a row on a fair coin is unlikely before it happens (about 3%), but after four heads the next flip is still 50/50. Our brains are pattern detectors tuned for a world where streaks usually mean something, but in pure chance they mean nothing at all.
Mean vs median: which should I use?
Use the mean for symmetric data with no extreme values, like well-behaved test scores. Use the median for skewed data like incomes or house prices, where a few huge values would drag the mean away from the typical case. When they disagree a lot, the data is skewed and the median is usually the more honest 'typical' value.
What does standard deviation actually tell me?
The typical distance of values from the mean, in the original units. A standard deviation of 5 points on a test means scores usually land within about 5 points of the average. Small means clustered and predictable; large means spread out and variable.
What is the difference between population and sample standard deviation?
Population standard deviation divides by N and describes the full group you measured. Sample standard deviation divides by N-1, which corrects the slight underestimate you get from a sample, and is the right choice when your data is a sample of something larger. Sample Size labels which one it uses.
What is variance?
The average of the squared distances from the mean. Squaring keeps above-and-below deviations from canceling out. The standard deviation is just the square root of variance, which puts the answer back into original units like dollars or points instead of squared dollars.
What is a percentile?
The value below which a given percent of the data falls. The 90th percentile of house prices is the price that 90% of houses cost less than. The 50th percentile is the median. It is a rank, not a percent score.
What is a z-score?
How many standard deviations a value sits from the mean. A z-score of +2 means the value is two standard deviations above average, which is roughly the 98th percentile in bell-shaped data. Z-scores let you compare values from completely different datasets on one scale.
What is the normal distribution?
The symmetric bell curve where most values cluster near the mean and extremes taper off. About 68% of values fall within one standard deviation, 95% within two. Many natural measurements approximate it, but plenty of real data (incomes, wait times) does not, so check the shape before assuming.
Does correlation imply causation?
No. Correlation says two things move together; causation says one makes the other happen. A hidden third factor often drives both, like summer driving both ice cream sales and drownings. Establishing causation needs controlled experiments or careful causal analysis, not just a correlation number.
How many data points do I need?
There is no universal minimum, but summaries from a handful of values are fragile: each new observation can swing them. As a rough guide, patterns worth trusting usually need dozens of observations, and formal inference wants enough data for the sampling noise to shrink below the effect you care about.
What is the mode good for?
It is the only 'average' that works on categories: the most common shirt size, the most frequent error code, the busiest hour. For numbers it is less informative than mean or median, but for categorical data it is often the single most useful summary.
How do I handle outliers?
First check whether they are errors (typos, wrong units) and fix or remove those. For genuine extremes, prefer the median and interquartile range over the mean and standard deviation, since the former barely budge when one value goes wild. Never delete outliers just to make the numbers look nicer.