Here are three sets of seven numbers each. The x values are the number of points scored by Kobe Bryant in the playoff games against the Houston Rockets, the y values are the points scored by the Lakers in those same games and the z values are the difference between the Lakers' score and the Rockets' score in each game.
x: __32__40__33__35__26__32__33
y: __92_111_108__87_118__80__89
z: __-8_+13_+14_-12_+30_-15_+19
1. Find the correlation coefficients for all three pairs of number sets, rx,y, rx,z and ry,z.
2. What are the cut-off values for 95% confidence and 99% confidence for correlation?
3. Which of the pairs of sets has the highest absolute correlation and what confidence level does that correlation exceed?
Round all answers to three digits. Answers in the comments.
Thursday, May 28, 2009
Wednesday, May 20, 2009
Topic for final exam
The final exam will be given at two times.
May 22: 8 a.m. to 10 a.m.
May 29: 10 a.m. to noon.
The final is comprehensive. You will need your yellow sheets, a calculator, scratch paper and a pencil.
The test will be four or five pages long. The amount from each part of the class will be
25%-30% from first exam
25%-30% from second exam
40-50% from after the second exam
On the list of topics below, any topic with an asterisk (*) means that though it might have been introduced before the first or second exam, it gets used throughout the class, so I don't necessarily count it as in the percentage of problem promised for each section.
Any topic in bold means you are expected to know how to get the answer without a formula or instructions being provided. In many cases, this means knowing how to use your calculator properly.
First exam
==========
frequency tables
stem and leaf plots
five number summary
box and whiskers
IQR and outliers for box and whiskers
Mean*, median*, mode, mid-range
parameter* and statistic*
population* and sample*
categorical data*
numerical data*
Bar charts
Pie charts
Line charts
Ogives
dotplots
percentage increase and decrease
contingency tables*
degrees of freedom*
conditional probability*
frequency and relative frequency
inclusion-exclusion
complementary event
order of operations*
Second exam through April 8
==========
standard deviation*
confidence intervals and margin of error
t-scores and z-scores*
raw scores, z-scores and percentages*
common critical values
Central Limit Theorem
Confidence of victory
After second exam
=======
Binomial coefficients and falling factorial
expected value of correct results
dependent and independent probabilities
Classic and modern parimutuel
expected value of a game
exactly r correct out of n trials
Bayesian probabilities
Hypothesis testing:
null hypothesis, alternative hypothesis, type I error, type II error
test statistic
threshold for xx% confidence (one-tailed high, one-tailed low, two-tailed)
one sample testing
two samples testing
Correlation (rx,y and the equation of the line yp = ax + b)
If you have any specific questions or want to make time to talk to me before the final, send me an e-mail and we can make an appointment.
May 22: 8 a.m. to 10 a.m.
May 29: 10 a.m. to noon.
The final is comprehensive. You will need your yellow sheets, a calculator, scratch paper and a pencil.
The test will be four or five pages long. The amount from each part of the class will be
25%-30% from first exam
25%-30% from second exam
40-50% from after the second exam
On the list of topics below, any topic with an asterisk (*) means that though it might have been introduced before the first or second exam, it gets used throughout the class, so I don't necessarily count it as in the percentage of problem promised for each section.
Any topic in bold means you are expected to know how to get the answer without a formula or instructions being provided. In many cases, this means knowing how to use your calculator properly.
First exam
==========
frequency tables
stem and leaf plots
five number summary
box and whiskers
IQR and outliers for box and whiskers
Mean*, median*, mode, mid-range
parameter* and statistic*
population* and sample*
categorical data*
numerical data*
Bar charts
Pie charts
Line charts
Ogives
dotplots
percentage increase and decrease
contingency tables*
degrees of freedom*
conditional probability*
frequency and relative frequency
inclusion-exclusion
complementary event
order of operations*
Second exam through April 8
==========
standard deviation*
confidence intervals and margin of error
t-scores and z-scores*
raw scores, z-scores and percentages*
common critical values
Central Limit Theorem
Confidence of victory
After second exam
=======
Binomial coefficients and falling factorial
expected value of correct results
dependent and independent probabilities
Classic and modern parimutuel
expected value of a game
exactly r correct out of n trials
Bayesian probabilities
Hypothesis testing:
null hypothesis, alternative hypothesis, type I error, type II error
test statistic
threshold for xx% confidence (one-tailed high, one-tailed low, two-tailed)
one sample testing
two samples testing
Correlation (rx,y and the equation of the line yp = ax + b)
If you have any specific questions or want to make time to talk to me before the final, send me an e-mail and we can make an appointment.
Monday, May 18, 2009
Class notes for 5/18: Regression and correlation
When we have a data set, sometimes we collect more than one variable of information about the units. For example, in our class survey, among the numerical variables were the height in inches, the GPA, the opinion about the difficulty of the class, age and average hours of sleep per night.
A question about two variables is if they are related to one another in some simple way. One simple way is correlation, which can be positive or negative. Here is a general definition of each.
Positive correlation between two numerical variables, call them x and y, means that the high values of x tend to be paired with the high values of y, the middle values of x tend to be paired with the middle values of y and the low values of x tend to be paired with the low values of y.
The variables x and y show negative correlation if that the high values of x tend to be paired with the low values of y, the middle values of x tend to be paired with the middle values of y and the low values of x tend to be paired with the high values of y.
If we pick two variables at random, we do not expect to see correlation. We can write this as a null hypothesis, where the test statistic is rx,y, the correlation coefficient. The sign of low correlation is rx,y = 0. The values of rx,y are always between -1, which means perfect negative correlation, and +1, which means perfect positive correlation.
The seventh page of the yellow sheets gives us threshold numbers for the 99% confidence level and 95% confidence level for correlation given the number of points n. For instance, when n = 5, the thresholds are .878 for 95% confidence and .959 for 99% confidence. This splits up the numbers from -1 to 1 into five regions.
-1 <= rx,y <= -.959: Very strong negative correlation
-.959 < rx,y <= -.878: Strong negative correlation
-.878 < rx,y < .878: The correlation is not particularly strong, regardless of positive or negative.
.878 <= rx,y < .959: strong positive correlation
.959 <= rx,y <= 1: very strong positive correlation
Just like with any hypothesis test, we should decide the confidence level before testing. This is a two-tailed test, because whether correlation is positive or negative, the relationships between number sets can often give us vital scientific information.
There is an important warning: Correlation is not causation. Just because two number sets have a relation, it doesn't mean that x causes y or y causes x. Sometimes there is a hidden third factor that is the cause of both of the things we are looking at. Sometimes, it's random chance and there is no causative agent at all.

Here is a set of five points, listed as (x,y) in each case.
(1,1)
(2,2)
(3,4)
(4,4)
(6,5)
As we can see, the points are ordered from low to high in both coordinates, so we expect some correlation. If we input the points into our calculator, we get a value for r (which is the same as rx,y) of .933338696..., which is strong positive correlation, but not very strong positive correlation. Assuming the 95% confidence level is good enough for us, we can use the a and b variables from out calculator to give us the equation of the line
yp = .797x + .649
This is called the predictor line (that's where the p comes from) or the line of regression or the line of least squares. Any such line for a given data set has two important criteria it meets. It passes through the centroid (x-bar, y-bar), the center point of all the data, and it minimizes the sum of the absolute values of the residuals, which is |y - yp| for all points.
Let's find the absolute values of the residuals for each of the five points, using the rounded values of a and b.
Point (1,1): |1 - .797*1 - .649| = 0.446
Point (2,2): |2 - .797*2 - .649| = 0.243
Point (3,4): |4 - .797*3 - .649| = 0.96
Point (4,4): |4 - .797*4 - .649| = 0.163
Point (6,5): |5 - .797*6 - .649| = 0.431
As we can see, the point (3,4) is farthest from the line, while the point (4,4) is the closest. The centroid (3.2, 3.2) is exactly on the line if you use the un-rounded values of a and b, and even using the rounded values, the centroid only misses the line by .0006.
In class, we used five points, but the last point was (5,6) instead of (6,5). This changes the numbers. rx,y goes up to .973328527..., which is above the 99% confidence threshold. The formula for the new predictor line is
yp = 1.2x - .2
Where we see the difference in these two different examples is in the residuals.
Point (1,1): |1 - 1.2*1 + .2| = 0
Point (2,2): |2 - 1.2*2 + .2| = 0.2
Point (3,4): |4 - 1.2*3 + .2| = 0.6
Point (4,4): |4 - 1.2*4 + .2| = 0.6
Point (5,6): |6 - 1.2*5 + .2| = 0.2
The closest point is now exactly on the line, which is a rarity, but even the farthest away point is only .6 units away, closer than the farthest away on the line with the lower correlation coefficient.
As we get more points in our data set, we lower our threshold that shows correlation strength. This way, a few points that are outliers do not completely ruin the chances of the data showing correlation, though sometimes strong outliers can mess up the data set and the correlation coefficient gets so close to zero that we cannot reject the null hypothesis that the two variables are not simply related.
A question about two variables is if they are related to one another in some simple way. One simple way is correlation, which can be positive or negative. Here is a general definition of each.
Positive correlation between two numerical variables, call them x and y, means that the high values of x tend to be paired with the high values of y, the middle values of x tend to be paired with the middle values of y and the low values of x tend to be paired with the low values of y.
The variables x and y show negative correlation if that the high values of x tend to be paired with the low values of y, the middle values of x tend to be paired with the middle values of y and the low values of x tend to be paired with the high values of y.
If we pick two variables at random, we do not expect to see correlation. We can write this as a null hypothesis, where the test statistic is rx,y, the correlation coefficient. The sign of low correlation is rx,y = 0. The values of rx,y are always between -1, which means perfect negative correlation, and +1, which means perfect positive correlation.
The seventh page of the yellow sheets gives us threshold numbers for the 99% confidence level and 95% confidence level for correlation given the number of points n. For instance, when n = 5, the thresholds are .878 for 95% confidence and .959 for 99% confidence. This splits up the numbers from -1 to 1 into five regions.
-1 <= rx,y <= -.959: Very strong negative correlation
-.959 < rx,y <= -.878: Strong negative correlation
-.878 < rx,y < .878: The correlation is not particularly strong, regardless of positive or negative.
.878 <= rx,y < .959: strong positive correlation
.959 <= rx,y <= 1: very strong positive correlation
Just like with any hypothesis test, we should decide the confidence level before testing. This is a two-tailed test, because whether correlation is positive or negative, the relationships between number sets can often give us vital scientific information.
There is an important warning: Correlation is not causation. Just because two number sets have a relation, it doesn't mean that x causes y or y causes x. Sometimes there is a hidden third factor that is the cause of both of the things we are looking at. Sometimes, it's random chance and there is no causative agent at all.

Here is a set of five points, listed as (x,y) in each case.
(1,1)
(2,2)
(3,4)
(4,4)
(6,5)
As we can see, the points are ordered from low to high in both coordinates, so we expect some correlation. If we input the points into our calculator, we get a value for r (which is the same as rx,y) of .933338696..., which is strong positive correlation, but not very strong positive correlation. Assuming the 95% confidence level is good enough for us, we can use the a and b variables from out calculator to give us the equation of the line
yp = .797x + .649
This is called the predictor line (that's where the p comes from) or the line of regression or the line of least squares. Any such line for a given data set has two important criteria it meets. It passes through the centroid (x-bar, y-bar), the center point of all the data, and it minimizes the sum of the absolute values of the residuals, which is |y - yp| for all points.
Let's find the absolute values of the residuals for each of the five points, using the rounded values of a and b.
Point (1,1): |1 - .797*1 - .649| = 0.446
Point (2,2): |2 - .797*2 - .649| = 0.243
Point (3,4): |4 - .797*3 - .649| = 0.96
Point (4,4): |4 - .797*4 - .649| = 0.163
Point (6,5): |5 - .797*6 - .649| = 0.431
As we can see, the point (3,4) is farthest from the line, while the point (4,4) is the closest. The centroid (3.2, 3.2) is exactly on the line if you use the un-rounded values of a and b, and even using the rounded values, the centroid only misses the line by .0006.
In class, we used five points, but the last point was (5,6) instead of (6,5). This changes the numbers. rx,y goes up to .973328527..., which is above the 99% confidence threshold. The formula for the new predictor line is
yp = 1.2x - .2
Where we see the difference in these two different examples is in the residuals.
Point (1,1): |1 - 1.2*1 + .2| = 0
Point (2,2): |2 - 1.2*2 + .2| = 0.2
Point (3,4): |4 - 1.2*3 + .2| = 0.6
Point (4,4): |4 - 1.2*4 + .2| = 0.6
Point (5,6): |6 - 1.2*5 + .2| = 0.2
The closest point is now exactly on the line, which is a rarity, but even the farthest away point is only .6 units away, closer than the farthest away on the line with the lower correlation coefficient.
As we get more points in our data set, we lower our threshold that shows correlation strength. This way, a few points that are outliers do not completely ruin the chances of the data showing correlation, though sometimes strong outliers can mess up the data set and the correlation coefficient gets so close to zero that we cannot reject the null hypothesis that the two variables are not simply related.
Labels:
class notes,
correlation,
predictor line,
regression
inputting the data for the worksheet using the TI-30XIIs
To get into the correct mode, follow these instructions.
[2nd][data]
move the underline so it is under 2-var, then press [enter].
[2nd][data]
move the underline so it so under clrdata, then press [enter].
now we enter the data.
[data]
x1 = 856.7[down]
y1 = 15.22[down]
x2 = 907.9[down]
y2 = 16.79[down]
x3 = 974.1[down]
y3 = 19.81[down]
x4 = 930.9[down]
y4 = 17.88[down]
x5 = 885.9[down]
y5 = 16.84[down]
x6 = 886.1[down]
y6 = 16.87[down]
x7 = 926.8[down]
y7 = 17.48[down]
x8 = 928.4[down]
y8 = 17.34[down]
x9 = 802.9[down]
y9 = 12.22[down]
x10 = 834.8[down]
y10 = 11.16[down]
x11 = 734.3[down]
y11 = 9.37[down]
x12 = 816.3[down]
y12 = 10.26[down]
x13=[statvar]
The numbers you need are
rx,y = .946053079... ~= .946 (Above the 99% threshold of .708)
a = 0.048098608... ~= .0481
b = -26.92322636... ~= -26.9232
So the formula to find the points on the line is yp = ax + b if you use the exact values and if using the approximation, yp = .0481x - 26.9232.
Answer for both versions to the nearest penny in the comments.
[2nd][data]
move the underline so it is under 2-var, then press [enter].
[2nd][data]
move the underline so it so under clrdata, then press [enter].
now we enter the data.
[data]
x1 = 856.7[down]
y1 = 15.22[down]
x2 = 907.9[down]
y2 = 16.79[down]
x3 = 974.1[down]
y3 = 19.81[down]
x4 = 930.9[down]
y4 = 17.88[down]
x5 = 885.9[down]
y5 = 16.84[down]
x6 = 886.1[down]
y6 = 16.87[down]
x7 = 926.8[down]
y7 = 17.48[down]
x8 = 928.4[down]
y8 = 17.34[down]
x9 = 802.9[down]
y9 = 12.22[down]
x10 = 834.8[down]
y10 = 11.16[down]
x11 = 734.3[down]
y11 = 9.37[down]
x12 = 816.3[down]
y12 = 10.26[down]
x13=[statvar]
The numbers you need are
rx,y = .946053079... ~= .946 (Above the 99% threshold of .708)
a = 0.048098608... ~= .0481
b = -26.92322636... ~= -26.9232
So the formula to find the points on the line is yp = ax + b if you use the exact values and if using the approximation, yp = .0481x - 26.9232.
Answer for both versions to the nearest penny in the comments.
Labels:
calculators,
correlation,
two variable statistics
Thursday, May 14, 2009
Class notes for 5/13: Two population tests for proportions and averages
The first hypothesis tests we studied were checking to see if an experimental sample produced a value that was significantly different from some known value produced either by math or by earlier experiments.
For example, in the lady tasting tea, since she has two choices each time a mixture is given to her, the math would say that her chance of getting it right by just guessing is 50% or H0: p = .5. In testing psychic abilities, there are five different symbols on the cards, so random guessing should get the right answer 1 out of 5 times, or 20%, so H0: p = .2.
In a test for average human body temperature, the assumption of 98.6 degrees Fahrenheit being the average came from an experiment performed in the 19th Century.
We can also do tests by taking samples from two different populations. The null hypothesis, as always, is an equality, the assumption that the parameters from the two different populations are the same. As always, we need convincing evidence that the difference is significant to reject the null hypothesis, and we can choose just how convincing that evidence must be by setting the confidence level, which is usually either 90% or 95% or 99%.
Two proportions from two populations

Like with the one proportion test, the test statistic is a z-score. We have the proportions from the two samples, p-hat1 = f1/n1 and p-hat2 = f2/n2, but we also need to create the pooled proportion p-bar = (f1 + f2)/(n1 + n2).
Here's an example from the polling data from last year.
Question: Was John McCain's popularity in Iowa significantly different from his popularity in Pennsylvania?
Let's assume we don't know either way, so it will be a two tailed test. Polling data traditionally uses the 95% confidence level, so that means the z-score will have to be either greater than or equal to 1.96 or less than or equal to -1.96 for us to reject the null hypothesis. Here are our numbers, with Iowa as the first data set.
f1 = 263
n1 = 658
p-hat1 = .400
f2 = 283
n2 = 657
p-hat2 = .430
p-bar = (263+283)/(658+657) = .415 (q-bar = .585)
Type this into your calculator.
(.400-.430)/sqrt(.415x.585/658+.415x.585/657[enter]
The answer is -1.103..., which rounds to -1.10. This would say the difference we see in the two samples is not enough to convince us of a significant difference in popularity for McCain between the two states, so we would fail to reject the null hypothesis. In the actual election, McCain had 45.2% of the vote in Pennsylvania and 44.8% of the vote in Iowa, which are fairly close to equal.
Two averages from two populations

In the tests to see if the average of some numerical value is significantly different when comparing two populations, we need the averages, standard deviations and sizes of both populations. The score we use is a t-score and the degrees of freedom is the smaller of the two sample sizes minus 1.
Question: Do female Laney students sleep more hours each night than male Laney students?
We will take our data from the larger of the two class surveys, Data Set #2. Here are the numbers for the students who submitted data, with the females listed as group #1. Again, let's assume a two-tailed test, since we don't have any information going in which should be greater, and let's do this test to 90% level of confidence.
H0: mu1 = mu2 (average hours of sleep are the same for males and females at Laney)
x-bar1 = 7.31
s1 = .94
n1 = 26
x-bar2 = 7.54
s2 = 1.47
n2 = 12
The degrees of freedom will be 12-1=11, and 10% in two tails gives us the thresholds of +/-1.796. Here is what to type into the calculator.
(7.31-7.54)/sqrt(.94^2/26+1.47^2/12)[enter]
-0.4971...
This number is between the thresholds, and so does not impress us enough to make us reject the null hypothesis. It's possible that larger samples would give us numbers that would show a difference, which if true would mean this example produced a Type II error, but we have no proof of that.
For example, in the lady tasting tea, since she has two choices each time a mixture is given to her, the math would say that her chance of getting it right by just guessing is 50% or H0: p = .5. In testing psychic abilities, there are five different symbols on the cards, so random guessing should get the right answer 1 out of 5 times, or 20%, so H0: p = .2.
In a test for average human body temperature, the assumption of 98.6 degrees Fahrenheit being the average came from an experiment performed in the 19th Century.
We can also do tests by taking samples from two different populations. The null hypothesis, as always, is an equality, the assumption that the parameters from the two different populations are the same. As always, we need convincing evidence that the difference is significant to reject the null hypothesis, and we can choose just how convincing that evidence must be by setting the confidence level, which is usually either 90% or 95% or 99%.
Two proportions from two populations

Like with the one proportion test, the test statistic is a z-score. We have the proportions from the two samples, p-hat1 = f1/n1 and p-hat2 = f2/n2, but we also need to create the pooled proportion p-bar = (f1 + f2)/(n1 + n2).
Here's an example from the polling data from last year.
Question: Was John McCain's popularity in Iowa significantly different from his popularity in Pennsylvania?
Let's assume we don't know either way, so it will be a two tailed test. Polling data traditionally uses the 95% confidence level, so that means the z-score will have to be either greater than or equal to 1.96 or less than or equal to -1.96 for us to reject the null hypothesis. Here are our numbers, with Iowa as the first data set.
f1 = 263
n1 = 658
p-hat1 = .400
f2 = 283
n2 = 657
p-hat2 = .430
p-bar = (263+283)/(658+657) = .415 (q-bar = .585)
Type this into your calculator.
(.400-.430)/sqrt(.415x.585/658+.415x.585/657[enter]
The answer is -1.103..., which rounds to -1.10. This would say the difference we see in the two samples is not enough to convince us of a significant difference in popularity for McCain between the two states, so we would fail to reject the null hypothesis. In the actual election, McCain had 45.2% of the vote in Pennsylvania and 44.8% of the vote in Iowa, which are fairly close to equal.
Two averages from two populations

In the tests to see if the average of some numerical value is significantly different when comparing two populations, we need the averages, standard deviations and sizes of both populations. The score we use is a t-score and the degrees of freedom is the smaller of the two sample sizes minus 1.
Question: Do female Laney students sleep more hours each night than male Laney students?
We will take our data from the larger of the two class surveys, Data Set #2. Here are the numbers for the students who submitted data, with the females listed as group #1. Again, let's assume a two-tailed test, since we don't have any information going in which should be greater, and let's do this test to 90% level of confidence.
H0: mu1 = mu2 (average hours of sleep are the same for males and females at Laney)
x-bar1 = 7.31
s1 = .94
n1 = 26
x-bar2 = 7.54
s2 = 1.47
n2 = 12
The degrees of freedom will be 12-1=11, and 10% in two tails gives us the thresholds of +/-1.796. Here is what to type into the calculator.
(7.31-7.54)/sqrt(.94^2/26+1.47^2/12)[enter]
-0.4971...
This number is between the thresholds, and so does not impress us enough to make us reject the null hypothesis. It's possible that larger samples would give us numbers that would show a difference, which if true would mean this example produced a Type II error, but we have no proof of that.
Tuesday, May 12, 2009
Class notes for 5/11: Hypothesis testing for the mean of a population
t-scores and p values
If we have a z-score between -3.5 and +3.5, Table A-2 lets us find the p value associated with that z-score accurate to four decimal places. For example, if z = 1.71, the p value is .9564, which is to say that z-score is higher than 95.64% of data in a normally distributed set.
The t-score table is smaller, and to read a t-score correctly, we also need n, the size of the sample, because that gives us the Degrees of Freedom, which is n-1 in this case.
If n=10, then d.f. = 9, and the t-score table reads as follows.
___________________________Area in One Tail_____________
_______0.005______0.01______0.025______0.05______0.10__________df=9___3.250_____2.821______2.262_____1.833_____1.383__________
If t = 1.71, that value isn't on our table, but because it lies between the values associated with 0.05 and 0.10, that means that score is in the top 10% of scores, but not in the top 5%.
If instead n=30 and d.f. = 29, here are the t-score values.
___________________________Area in One Tail_____________
_______0.005______0.01______0.025______0.05______0.10__________df=29__2.756_____2.462______2.045_____1.699_____1.311__________
Now a t-score of 1.71 lies between 2.045 and 1.699, which means it is in the top 5%, but not the top 2.5%.
Like the z-score table, the t-score table is symmetric about the value t=0. If d.f.=29, t=-1.71 is a score in the bottom 5%, but not the bottom 2.5%.
Hypothesis testing for the mean of a population

Hypothesis testing for the mean of a population assumes we know the population mean from some previously obtained information. Perhaps that mean has changed over time or the previous information wasn't correct to begin with, but the null hypothesis assumes we know that mean, which we call mux. If we take a sample from the population, we will get the values x-bar, sx and n, and using those values and mux, we can get the t-score.
Just like with the hypothesis test for a proportion, the test can be one-tailed high, one-tailed low or two-tailed.
For example, if we were testing people who had studied using a special method and we were checking scores on a standardized test, we would only be impressed if the average went up, so a one-tailed high test would be appropriate.
If the experiment was dealing with a cholesterol drug, we would want to see a lower average reading, and a one-tailed low test would be used.
If we assume the average duration of a pop song on the charts today is the same as the duration of pop songs in the seventies, we can't assume beforehand if the new readings will be higher or lower, and would be surprised if the new average were significantly different in either direction, so a two-tailed test would be appropriate.
Here is some data we went over in class. In most textbooks, the 'normal' human body temperature is listed at 98.6 degrees Fahrenheit, based on the work of Dr. Carl Wunderlich back in the 19th Century. If we do a test, it should be a two-tailed test, since we would be surprised if the normal temperature is significantly higher or significantly lower than this. Since this is a medical experiment, let's use the 99% level of confidence.
The size of the sample was n=103, which means the degrees of freedom are 102. Our table doesn't have a listing for d.f.=102, and the next lowest available value is d.f.=100. Here are the table values for that row of Table A-3.
___________________________Area in Two Tails____________
_________0.01______0.02_______0.05______0.10______0.20__________df=100__2.626_____2.364______1.984_____1.660_____1.290__________
With a two tailed test at the 99% confidence level, this means we want the 0.01 column. The "middle" 99% of data lies between t-scores of -2.616 and 2.616. If the t-score lies in that range, we will fail to reject H0. If it is greater or equal to 2.616 or less than or equal to -2.616, we will reject H0.
The values from this study found that x-bar = 98.2 degrees and the standard deviation sx was 0.62. Plugging into our t-score equation from above, we get
t = (98.2-98.6)/0.62*sqrt(103) = -6.547671977... ~= -6.548.
We don't get an exact p value for a number so far away from zero, but if we look at outlier z-score table, we can roughly approximate that this p value is somewhere around 1 in 1,000,000,000. We can say with 99% confidence that the average body temperature is not 98.6 degrees, but probably close to the sample average of 98.2 degrees. Our p value shows we could qualify for even greater confidence with our statement, but very rarely do tests ask for more than 99% confidence, and changing the criteria after the fact is not recommended. Still, publishing this incredibly tiny p value will convince people who can read a statistical report that the evidence is very strong indeed.
This test also changed the idea of what should constitute a fever. Instead of one temperature of 100.4 degrees Fahrenheit being the absolute gauge, the temperature will fluctuate depending on the time of day, as do the normal temperature readings.
If we have a z-score between -3.5 and +3.5, Table A-2 lets us find the p value associated with that z-score accurate to four decimal places. For example, if z = 1.71, the p value is .9564, which is to say that z-score is higher than 95.64% of data in a normally distributed set.
The t-score table is smaller, and to read a t-score correctly, we also need n, the size of the sample, because that gives us the Degrees of Freedom, which is n-1 in this case.
If n=10, then d.f. = 9, and the t-score table reads as follows.
___________________________Area in One Tail_____________
_______0.005______0.01______0.025______0.05______0.10__________df=9___3.250_____2.821______2.262_____1.833_____1.383__________
If t = 1.71, that value isn't on our table, but because it lies between the values associated with 0.05 and 0.10, that means that score is in the top 10% of scores, but not in the top 5%.
If instead n=30 and d.f. = 29, here are the t-score values.
___________________________Area in One Tail_____________
_______0.005______0.01______0.025______0.05______0.10__________df=29__2.756_____2.462______2.045_____1.699_____1.311__________
Now a t-score of 1.71 lies between 2.045 and 1.699, which means it is in the top 5%, but not the top 2.5%.
Like the z-score table, the t-score table is symmetric about the value t=0. If d.f.=29, t=-1.71 is a score in the bottom 5%, but not the bottom 2.5%.
Hypothesis testing for the mean of a population

Hypothesis testing for the mean of a population assumes we know the population mean from some previously obtained information. Perhaps that mean has changed over time or the previous information wasn't correct to begin with, but the null hypothesis assumes we know that mean, which we call mux. If we take a sample from the population, we will get the values x-bar, sx and n, and using those values and mux, we can get the t-score.
Just like with the hypothesis test for a proportion, the test can be one-tailed high, one-tailed low or two-tailed.
For example, if we were testing people who had studied using a special method and we were checking scores on a standardized test, we would only be impressed if the average went up, so a one-tailed high test would be appropriate.
If the experiment was dealing with a cholesterol drug, we would want to see a lower average reading, and a one-tailed low test would be used.
If we assume the average duration of a pop song on the charts today is the same as the duration of pop songs in the seventies, we can't assume beforehand if the new readings will be higher or lower, and would be surprised if the new average were significantly different in either direction, so a two-tailed test would be appropriate.
Here is some data we went over in class. In most textbooks, the 'normal' human body temperature is listed at 98.6 degrees Fahrenheit, based on the work of Dr. Carl Wunderlich back in the 19th Century. If we do a test, it should be a two-tailed test, since we would be surprised if the normal temperature is significantly higher or significantly lower than this. Since this is a medical experiment, let's use the 99% level of confidence.
The size of the sample was n=103, which means the degrees of freedom are 102. Our table doesn't have a listing for d.f.=102, and the next lowest available value is d.f.=100. Here are the table values for that row of Table A-3.
___________________________Area in Two Tails____________
_________0.01______0.02_______0.05______0.10______0.20__________df=100__2.626_____2.364______1.984_____1.660_____1.290__________
With a two tailed test at the 99% confidence level, this means we want the 0.01 column. The "middle" 99% of data lies between t-scores of -2.616 and 2.616. If the t-score lies in that range, we will fail to reject H0. If it is greater or equal to 2.616 or less than or equal to -2.616, we will reject H0.
The values from this study found that x-bar = 98.2 degrees and the standard deviation sx was 0.62. Plugging into our t-score equation from above, we get
t = (98.2-98.6)/0.62*sqrt(103) = -6.547671977... ~= -6.548.
We don't get an exact p value for a number so far away from zero, but if we look at outlier z-score table, we can roughly approximate that this p value is somewhere around 1 in 1,000,000,000. We can say with 99% confidence that the average body temperature is not 98.6 degrees, but probably close to the sample average of 98.2 degrees. Our p value shows we could qualify for even greater confidence with our statement, but very rarely do tests ask for more than 99% confidence, and changing the criteria after the fact is not recommended. Still, publishing this incredibly tiny p value will convince people who can read a statistical report that the evidence is very strong indeed.
This test also changed the idea of what should constitute a fever. Instead of one temperature of 100.4 degrees Fahrenheit being the absolute gauge, the temperature will fluctuate depending on the time of day, as do the normal temperature readings.
Labels:
hypothesis testing,
p-values,
Student's t-scores
Monday, May 11, 2009
Practice true-false questions about hypothesis testing.
1. If the confidence level is 90%, it is more common to make Type I errors than it is with a confidence level of 99%.
2. Proportion tests are never two-tailed.
3. If we reject H0, but in reality we shouldn't have, we have made a Type II error.
4. If we have a z-score of -1.38, we would reject H0 in a 90% confidence one-tailed low test.
5. If we have a t-score of -1.38 and n=7, we would reject H0 in a 90% confidence one-tailed low test.
6. The null hypothesis is always stated as an equation.
7. You can never start an experiment by assuming the alternate hypothesis is true.
8. If we have a z-score of -1.68, we would reject H0 in a 90% confidence two-tailed test.
9. For us to be 95% confident the lady tasting tea knew what she was doing, she had to get at least 95% of her answers correct.
10. You are allowed to do an experiment and decide what confidence level you want to use after you seen the results.
Answers in the comments.
2. Proportion tests are never two-tailed.
3. If we reject H0, but in reality we shouldn't have, we have made a Type II error.
4. If we have a z-score of -1.38, we would reject H0 in a 90% confidence one-tailed low test.
5. If we have a t-score of -1.38 and n=7, we would reject H0 in a 90% confidence one-tailed low test.
6. The null hypothesis is always stated as an equation.
7. You can never start an experiment by assuming the alternate hypothesis is true.
8. If we have a z-score of -1.68, we would reject H0 in a 90% confidence two-tailed test.
9. For us to be 95% confident the lady tasting tea knew what she was doing, she had to get at least 95% of her answers correct.
10. You are allowed to do an experiment and decide what confidence level you want to use after you seen the results.
Answers in the comments.
Subscribe to:
Posts (Atom)
