Tuesday, January 21, 2014

Lecture notes for 21 Jan. 2014

Variables and values


Each element of a data set has different pieces of information that are collected. For example, let's say this some of the information on a driver's license.

Height: 5'11"
Weight: 175 lbs.
Eyes: Brown
Hair: Black
Date of birth: 7/27/1985

The variables here are Height, Weight, Eyes, Hair and Date of birth. Height, Weight and Date of birth are numerical variables, since the answers are numbers. Eyes and hair are categorical variables, where the answers are not numbers. "175 lbs." is the value associated with the variable Weight.

Continuous variables versus discrete variables

Consider shoe sizes. In the United States, the sizes are ½ apart, so the count goes 5, 5½, 6, 6½, 7, 7½, 8, 8½, etc. What this means is if you give me a show size, I can tell you then next highest and the next lowest. According to Wikipedia, the smallest shoe size is a 1, so for that size there is no next lowest, but the next highest is 1½. A situation like this means the variable is discrete.

Just because two people wear the same size shoe does not mean they have the same size feet. For example, if a man's foot is between 10.92" and 11.08" long, the most comfortable shoe should be a size 9. You cannot say that there is a size of foot that is "the next size up" or "the next size down" from any given foot. No matter how close two feet are in size, if they are not exactly the same, it should be possible to find a foot that is in between the two that are chosen. When that is the case, when there is no defined "next size" either up or down, we say a variable is continuous.

Types of numerical data 
Coded numerical or nominal data: Usually, with numbers there is a meaning we can give to the ideas of "more" and "less". With coded data, we don't necessarily have that. Examples are zip codes, social security numbers and driver's licenses. Finding the average zip code of a group of people is meaningless, as is the mid-range and median. Finding the mode means that is the most popular of the possible zip codes, and that does give valuable information.

Ordinal data: Here, the idea of a > b has meaning, but the distance between units isn't the same. In a ranking system, it's better to be first than it is to be second, but we can't say how much better, and we don't know if the difference between first and second is the same as the distance between second and third.

Often, when we switch from an ordered categorical system like grades (A, B, C, D, F) to the numbers used for grade points (4.0, 3.0, 2.0, 1.0, 0.0), the choice of what numbers to use is arbitrary. Is the distance from an A to B really the same as the distance from a C to a D? Is getting an A in one class and a C in another really the same as getting two Bs, since both would be a 3.0 Grade Point Average (GPA). How about 2 As and a D, which is 3.0, or 3 As and an F? Should all those situations be counted the same way?

The feeling of this instructor is that it should not. Again, like coded data, ordinal data can use some of the measures of center, like median and mode, but average and mid-range do not give useful information.

Interval data: This is the minimum requirement need for mean to make sense, the idea that the distance between two numbers, a - b, has a consistent meaning, like degrees in temperature readings or the number of strokes taken to complete a round of golf. In these system, the number zero does not mean the complete absence of a thing, so it dividing one number by another from the data set doesn't give meaningful information, but taking and average is about adding values together and dividing by the number of values, so an average temperature or an average of the scores in four rounds of golf does produce a useful statistic.

Rational data: This is data where not only a - b means something, but also a/b. The difference between interval and rational data is the meaning of the number zero. If zero indicates the complete lack of a thing, then we can talk about something between twice as much as another thing, or 10% less. A lot of numerical systems of measurement are rational, but not all.


Measures of center
 
Mode
Type of data: any data can be used, but only if there are duplicate values on the list.
Method: Find the most common value. If there is a tie for most common, there can be more than one mode.

Median
Type of data: numerical or ordered categorical
Method: Put the values in order and find the "middle value" which is the value in position (n+1)/2. If n is odd, there is a single median value on the list. If n is even, there are two middle values on the list, and if numerical, take the average of the two. If the data is categorical and the two values aren't the same, the median lies between two categories.

Mean (average)
Type of data: numerical
Method: add up all the numbers and divide by n (or N), the number of things on the list.

Mid-range
Type of data: numerical
Method: (high + low)/2

Let's take the Games Behind data from today's handout and find the mode, median, mean and mid-range. The dash at the top (-) actually signifies 0 games behind.

Data set: 0, 4½, 12, 13, 13, 13, 15½, 16½, 16½, 18½, 18½, 20, 20½, 22½, 26
size of data set: n = 15

Mode: Mode is easy when the data is put in order like this. There are three 13 values, but only two 16½ values and two 18½ values. The mode is 13.

Median: Since n = 15,  the middle position counting from left or right is (15+1)/2 = 16/2 = 8. There is a single number in the eighth position whether counting from the left or the right.

0, 4½, 12, 13, 13, 13, 15½, 16½, 16½, 18½, 18½, 20, 20½, 22½, 26

The first 16½ on the list is the median.

Mean: The total is 230, so the mean is 230/15 = 15.3333..., which we will round to the nearest tenth, so the mean is rounded to 15.3

Mid-range: The biggest number is 26, the smallest number is 0 and (26+0)/2 = 13, so the mid-range is 13.

Saturday, January 18, 2014

Syllabus for Spring 2014

Math 13: Introduction to Statistics - class code 22723
Spring 2014 – Laney College
Tuesday Thursday: 10:30 am-12:20 pm

Instructor: Matthew Hubbard 
Email address: mhubbard@peralta.edu
No recommended textbook
class website: http://budgetstats.blogspot.com/

Office hours: 
 T-Th: 1:30-1:55 pm G-201 (Math Lab)
T-Th: 6:30-6:55 pm G-201 (Math Lab)

Add and drop class dates
Last date to add: Sat., Feb. 1
Last date to drop class without a “W”: Sat., Feb. 1
Last date to drop class with a “W”: Sat., May 3

Holiday schedule for Tuesday-Thursday classes
Spring break April 14 to 19

Test dates:
Midterm 1: Thurs., Feb. 27
Midterm 2: Thurs., Apr. 10
Comprensive Final: Thursday, May 22 10:00 am to noon

Homework to be turned in: Assigned on Thursdays, due the next Tuesday.
Homework can be turned into Mr. Hubbard’s box in the Math Lab G-201, open Tuesday through Thursday from 9:00 am to 7:00 pm
OR can be turned in by e-mail to mhubbard@peralta.edu.

Late homework accepted AT THE BEGINNING of the next class (usually Thursday)

Quizzes: One every Thursday, except the first and last week and weeks with a midterm

Grading system
20% Homework
  5% Labs
25% Quizzes best 2 of 3
25% Midterm I best 2 of 3
25% Midterm II best 2 of 3
25% Final

Lowest two of the homework scores will be dropped from the total.
Lowest two of the quiz scores will be dropped from the total.
Lowest total out of 100 points the quiz total and two midterms will be dropped from the final grade.
Anyone getting a higher grade out of 100 points on the final than the weighted average of all grades combined will get the final percentage instead deciding the final grade. This option is only available to students who have missed at most two homework assignments.
There are no make-up quizzes. Midterms can be made up if the student gives prior notice and has time to take the test before the next class meeting.
Anyone whose class average is over 97% going into the final can skip the final and get an A in the class.

Class rules: All cell phones and electronic communication devices off during class.
No hats, hoodies or headphones worn during quizzes and exams.
No calculators that also combine a cell phone or text message machine.

Recommended calculator: TI-30XIIs (any calculator with at least two lines of output will do, the TI-30XIIs is the cheapest that does all the things you need to do in this class. If you need help with any Texas Instruments calculator, I should be able to steer you in the right direction. I haven’t used other brands of calculators as much.)

Academic honesty: All assignments you turn in, homework, exams and quizzes, must be your own work. Anyone caught cheating on these assignments will be punished, where the punishment can be as severe as getting a zero on the assignment.


Student learning outcomes

Math 13 — Introduction to Statistics
1. Describe numerical and categorical data using statistical terminology and notation.
2. Analyze and explain relationships between variables in a sample or a population.
3. Make inferences about populations based on data obtained from samples.
4. Given a particular statistical or probabilistic context, determine whether or not a particular analytical methodology is appropriate and explain why.

The reciprocal relationship

The teacher will be on time and prepared to teach the class.
The students will be on time and prepared to learn.

The teacher will present the material to the best of his ability.
The students will absorb the material to the best of their ability. They will ask questions when topics are not clear.

The teacher will do his best to answer the questions the students ask about the material, either by repeating an answer with more details included or by taking a different approach to the material that might be clearer to some students.
The students will understand if the teacher feels a topic has been covered enough for the majority of the class and will accept questions being answered outside the class, either in extra time or through written communication.

The teacher will do his best to keep the class about the material. Personal details and distractions that are not germane to the class should not be part of the class.
The students will do their best to keep the class about the material. Questions that are not about the topic should be avoided. Distractions like cell phones and texting are not welcome when the class is in session.

The teacher will give assignments that will help the students master the skills required to pass the course.
The students will put in their best efforts to complete the assignments.
When the assignments are completed, the teacher will make every effort to get the assignments graded and back to the students in a timely manner, by the next class session whenever possible.

The teacher will present real life situations where the skills being learned will be used when they exist. In math, sometimes a particular skill is needed in general to solve later problems that will have real life applications. Other skills have the application of “learning how to learn”, of committing an idea to memory so that committing other ideas to memory becomes easier in the long run.
The student has the right to ask “When will I use this?” when dealing with mathematical topics. Sometimes, the answer is “We need this skill for the next skill we will learn.” Other times, the answer is “We are learning how to learn.” Both of these answers are as valid in their way as “We will need this to understand perspective” or “We use this to balance our checkbooks” or “Ratios can be used to change a recipe that serves three people to one that serves ten people” or other real life applications.

Thursday, October 4, 2012

Confidence intervals for proportions.

In Excel, we can find the z-scores that give us the end points of the confidence intervals as follows.

90% confidence interval.  The middle 90% is between 5% and 95%. The Excel formulas are

=norminv(.05, 0, 1)

= norminv(.95, 0, 1)

Rounded to four places after the decimal, we get -1.645 and 1.645.

95% confidence interval.  The middle 95% is between 2.5% and 97.5%. The Excel formulas are

=norminv(.025, 0, 1)

= norminv(.975, 0, 1)

Rounded to four places after the decimal, we get -1.960 and 1.960.


99% confidence interval.  The middle 90% is between 0.5% and 99.5%. The Excel formulas are

=norminv(.005, 0, 1)

= norminv(.995, 0, 1)

Rounded to four places after the decimal, we get -2.576 and 2.576.
In our first sample of 180 m&ms, there were 15 red m&ms. This says p_hat(red) = 15/180 ~= 8.3% when rounded to the nearest tenth of a percent.

The standard deviation for this proportion sp_hat(red) = sqrt(.083(1-.083)/180) or .020600514, which I will round to 2.1%

Here are the confidence intervals for this sample.

90% confidence interval.

0.083 + 1.645(.021) = .117545 about 11.8%
0.083 - 1.645(.021) = .048455 about 4.8%

Given this sample, we are 90% confident the true proportion of red m&ms in the population is between 4.8% and 11.8%


95% confidence interval.

0.083 + 1.960(.021) = .12416 about 12.4%
0.083 - 1.960(.021) = .04184 about 4.2%

Given this sample, we are 95% confident the true proportion of red m&ms in the population is between 4.2% and 12.4%



99% confidence interval. 

0.083 + 2.576(.021) = .137096 about 13.7%
0.083 - 2.576(.021) = .028904 about 2.8%

Given this sample, we are 99% confident the true proportion of red m&ms in the population is between 2.8% and 13.7%


As we increase the confidence, the interval gets larger.

As the sample size n gets larger, sqrt(p_hat*q_hat/n) will tend to get smaller, so larger sample sizes will give us smaller intervals, assuming p_hat doesn't change much.

Monday, September 3, 2012

Notes for 9/5/12


Here is a link to the topics for this Wednesday, September 5.

We will also discuss ordered and unordered categorical data and the problems with Excel using some of these ideas.

Wednesday, August 24, 2011

Early definitions.

A link to a post about definitions.  The post has more information than we went over in class.  We will get to this stuff all this stuff in the next few classes on Friday and Monday.

These notes are taken from a class taught in two hour sessions, so there will often be extra info.  You can read it to stay ahead.

Saturday, December 4, 2010

Hans Rosling 200 countries, 200 years, 4 minutes



Here is Hans Rosling's 200 countries and 200 years in four minutes. Here are some questions from the four minutes.

What does the x axis represent?
Are the x axis numbers linearly larger? (This question was addressed in class.)
What does the y axis represent?
Are the y axis numbers linearly larger? (This question was addressed in class.)
What does the color of a dot represent?
What does the size of a dot represent?
Which continent has the most countries getting healthier and wealthier in the 19th Century, due in large part to the Industrial Revolution?
Rosling stops for a pair of global catastrophes that overlapped in time. What are they?
At the end of World War II, what country is in the lead in terms of health and wealth?
In 2009, what country is in the lead in terms of health and wealth?
In 2009, what country is far behind in terms of health and wealth?
When splitting up China, Shanghai is about on par with ______ while rural parts of Guizhou are on par with _____.

Watch the video and answer the questions. The video doesn't fit the screen very well, so click on it and watch it on YouTube.
(Answers in the comments.)

Thursday, November 4, 2010

Stuff to review for the second midterm.

The second midterm will cover topics from Homeworks 6, 7, 8, 9 and 10. These will include:

Distributions from independent trials (sampling with replacement)
Distributions from dependent trials (sampling without replacement)
Margins of error from opinion poll percentages (also know as the 95% confidence interval)
The confidence of victory formula
Sentences that explain margin of error and confidence of victory numbers
Modern and classic pari-mutuel payoffs (profit and risk)
Expected Value of a win-lose game
Hypothesis testing
Rejecting the null hypothesis and failing to reject the null hypothesis
Type I error (rejecting the null when you shouldn't)
Type II error (failing to reject the null when you should)
Formulas for creating the test statistic for hypothesis testing (z-scores and t-scores)
Finding the threshold number for hypothesis testing (one-tailed high, one-tailed low, two tailed)