Get data from newspaper: marriage announcements. For example, you could go to
2. Enter their ages into the regression program at http://vassarstats.net/
3. For people getting married,
a. Is there a relationship between the man’s and woman’s ages?
b. What is the Pearson r between groom and bride ages?
What type of relationship does that r indicate—and is it a weak or strong relationship?
2. What does it mean if your analysis shows that p < .05? Answer this question twice. First, answer it as though you were talking to a scientist who knows statistical terminology. Then, answer is as though you were talking to someone who knew nothing about statistics.
4. Does the woman in this sample tend to be younger than the man they are marrying?
a. By how many years?
b. Assuming your sample was a random sample of men and women in the U.S., what statistical test would you use to find out whether American women tend to be significantly (reliably) younger than the men they marry?
5. Write an equation that predicts the woman’s age from the man’s age.
Your equation will look like this
Y(Bride’s age) = Slope * X (Man’s age) + b (the y-intercept)
You will fill in values for Slope and b.
Use that equation to predict the first bride’s age. That is, plug in the first groom’s age for X and see what “Y” (bride’s age) you calculate.
“Predicted” bride’s age _______
Actual bride’s age ________
How far off was the first bride’s predicted age from her actual age?
How close is this error in prediction to the number in the “Residuals” column of the first bride’s row?
First names may have more of an effect on people than you think. There is a relationship between names and where you live (a disproportionate number of Marys move to Maryland) and what you do (a disproportionate number of people named Dennis are dentists). Does first name also affect whom you marry? One way to start answering this question is to ask whether a greater proportion of couples both have J (your professor may assign you a different letter) as the first letter of their first name than would be expected by chance.
Step 1: Find out the proportion of couples who both have J as the first letter of their first names. (Example: If 6 couples had “A” as their first initial and there were 100 couples in your sample, the proportion of couples who both have “A” as their first initial would be [6/100].)
1. How many couples both had “J” as the first letter of their first names? _____
2. How many couples were in the sample? _____
3. What proportion of couples both had “J” as the first letter of their first names? ____
Step 2: Find out what proportion of couples we would expect—by chance alone—to share J as their first initial. That is, find out the proportion of couples whose first initials should have matched if people paired up randomly-- without regard to first names.
1a. In your sample, how many men’s first names start with J ? ____
1b. What is the number of men in your sample? ____
1c. The proportion of men whose first name starts with J
is ____ out of ______
2. The proportion of women whose first name starts with J is is ____ out of ______.
3. What would be the expected proportion of couples who would share their first name? (This is obtained by multiplying the proportion you got in “1c” by the proportion you got in “2.” The logic is the same as if you were trying to compute the odds of getting two heads in two flips: You multiply the chance of getting a heads on the first flip [.5] X the chances of getting a heads on the second flip [.5] to get .25). Similarly, the chances of getting “double sixes” in a throw of the dice are 1/6 X 1/6 = 1/36).
Step 3: See whether the percentage of couples who both have “J” as the first letter of their first names is significantly greater than would be expected by chance alone by
b. Clicking on “Proportions”
c. Clicking on “The Confidence Interval of a Proportion.”
d. Putting the number of couples whose first names had the same letter in the box for k =
e. Putting the total number of couples in the blank for n =
f. Pressing the “Calculate” button to obtain the 95% confidence interval for your obtained proportion.
g. State the lower and upper limits of the confidence interval.
h. The question now is whether your confidence interval for your obtained proportion excludes the proportion that would be expected to occur by chance alone. If people are more likely to marry others who share their first name, you would expect that the lower limit of the 95% confidence interval is higher than the expected proportion of couples who would share the letter J by chance (In Step 2, you calculated the expected proportion of couples who would share the same letter by chance). If the lower limit of the confidence interval based on your observed data is greater than what would be expected if only chance were at work, you can conclude that people with a first initial of J are more likely to marry other people with that same first initial. (The logic is similar to that of determining whether a coin is biased. If we flip a coin 100 times and it comes up heads 85 times, we suspect it is a biased coin because we know that a fair coin is unlikely to come up heads more than 80 times in a 100 flips.) Does the confidence interval include the expected proportion of couples who would share the letter by chance?
i. What is your conclusion?