click below
click below
Normal Size Small Size show me how
Biology Statistics
Unit 1
| Question | Answer |
|---|---|
| When asked to explain a line graph with a positive slope where the line suddenly plateaus... | |
| Mean | -average -sum up all the data points in the data set (∑X) and divide this # by the total # of data points (N) |
| Median | midpoint of the data -order from largest to smallest and choose the midpoint -not distorted by extreme values -in an even set of numbers, average the middle 2 #s |
| mode | -another measure of the average -the value that appears most often |
| bimodal distribution (opposite of unimodal) | -there are two clusters or ranges where data clusters most frequently --> two modes -these two modes don't have to be equal. it just has to have two distinct peaks in two separate data clusters |
| measures of variability | -describe the extent to which #s in a set diverge from the central tendency (how spread out data is around the center) -includes range, standard deviation, variance -a set has a greater variability if its values are further from the central tendency |
| range | the distance between the lowest and highest values in a data set |
| standard deviation | -how far on average any data point is from the mean WITHIN A SINGLE SAMPLE (amount of variation in a set of values) -a lower SD = closer scores are on average to mean -high SD = scores are widely spread out -square root of variance |
| variance (calculations) | 1. find mean 2. find how far each score is from the mean (score - mean) 3. square each difference 4. add the squared differences 5. if it's a population divide by N, if it's a sample divide by degrees of freedom |
| central tendency | -the value that represents the center or typical value of a data set -there are three common measures of central tendency: mean, median, mode -tells us where the center is -NOTE: a statistic is a broader term that encompasses central tendency |
| normal distribution curve | -bell curve: most values cluster around center - left & right sides are mirror images -mean = median = mode --> all three occur @ the center/peak - tails approach the x-axis but NEVER touch it -it can have differ. SDs -APPLITES TO CONTINUOUS DATA |
| degrees of freedom | n - 1 n =# if values in the set -how many pieces of info are free to vary after accounting for restrictions -the last value can't vary cauz once the other #s are chosen, it must be a specific # to make the set have the required mean |
| population | -size represented by N -an entire group you want to know about |
| sample | -size represented by n -a smaller group taken from a population that you can actually study |
| variance VS standard deviation | -both measure the same thing (how spread out the data are around the mean) -they have different units -VARIANCE: it's units are squared --> take original units of the data and square it -SD: original units of the data |
| one standard deviation from the mean | -the range of values that are within 1 SD of above or below the mean -on a normal distribution, about 68.2% of data falls within SD1 and 95% within 2SD |
| how to find 1SD | both add and subtract the standard deviation from the mean -this gives you the RANGE of values within 1SD -if you wanted 2SD, you would add and subtract the (standard deviation times 2) |
| standard error (definition) | SE reveals how much a statistic would vary if you repeatedly took new samples. (how much a statistic varies from sample to sample) -how precise an estimate is -measures spread of statistics across many diff. samples |
| standard error (how to calculate) | divide the standard deviation by the square root of the sample size |
| error bar (for standard error) | -a line drawn on a graph that shows the uncertainty or variability around an estimate -large SD = large error bar -add and subtract the SE to the mean: this reveals the range of the error bar --> mark both pts and draw a line between them |
| Why are smaller samples less accurate? | -they have a higher sensitivity to outliers -in larger sets, unusually high & unusually low value are more likely to balance each other out (improves precision) -Also, in the SE formula, you have to divide by n, so a larger n results in less error |
| NOTE: just given the means (or another measure of central tendency) of two data sets, you can't compare the two accurately.- You need to know the SD or variance as well as the sample size | |
| how to make a whisker plot | minimum: lowest # in set maximum: highest # center of the plot (marked w/ line in middle of the box) = median ORDER DATA & SPLIT IN 1/2: - lower quartile (Q1) = line below the box (median of lower 1/2) -upper quartile (Q3) (median of upper 1/2) |
| Why are normal distribution curves continuous | a normal curve assumes every value in the interval (within the spread) is possible. But, discrete data only includes distinct possible values w/ gaps between pts |
| when can a normal continuous distribution approximate discrete data sets | -discrete values are closely spaced -large n value (outcomes are more finely spaced relative to the overall spread of the distribution) - shape of discrete distribution must be roughly symmetric, concentrated around center, & tapering towards ends |
| Why are the tail ends of a normal distribution asymptotic to the x-axis (they approach it forever but never touch) | -extremely far out values have extraordinarily tiny probabilities, but the model never declares them completely impossible. -normal distribution is a mathematical model of reality, not reality itself. Real-world variables, however, have actual limits. |
| Histogram | graphical representation of numerical data that organizes data points into specified ranges. x: the value you are measuring (continuous ranges of data) y: the number of counts for a particular measurement (frequency of the mesurment) |
| standard error of the mean (SEx) | when a # of repeated samples are taken from a pop., a mean can be calculated for each sample. When plotted on a histogram, they are normally distributed -as the n of each sample increases, the curve becomes narrower (SD decreases and SE decreases) |
| how to estimate standard error of the mean | SE = sample SD/ square root of n NOTE: a histogram of samples means plots means rather than individual values, so its spread is measured by the standard error not the SD |
| the standard error of the mean INCREASES or DECREASES as sample size increases. | decreases |
| what does a larger n value mean for standard deviation of a sample? | as n gets larger, n-1 gets closer to N & s gets closer to the true population SD -in small sets, values are likely near average so you won't hit extreme values: smaller samples will look clumped together & larger samples will look more spread |
| The formula for estimating the standard error of the mean is only reliable in specific cenarios. | -it is only effective for large sample sizes -it includes SD in the formula, & the SD of small samples sizes underestimates the standard deviation of the population (it misses extreme values) |
| 95% confidence intervals | -way to determine if a sample represents the entire population -allows you to make a claim about the reliability of your data sample (large n increases reliability) - - |
| How to calculate the confidence interval | SE times t* -they should give you t |