Showing posts with label Statistics Basic Concepts. Show all posts
Showing posts with label Statistics Basic Concepts. Show all posts

Wednesday, June 20, 2012

No More Crying About Your Stats Class and Cry1Ab: An Application of the Coefficient of Variation


Lots of times students complain that either their statistics classes used silly examples that were too simple to ever be realistic, or that their course was too complicated and thus they leave the class without the capability of  any practical application.  A recent study looking at the safety of GMO corn provides a great case study for the practical application of the coefficient of variation (CV). 

In ‘Maternal and fetal exposure to pesticides associated to genetically modified Foods in Eastern Townships of Quebec, Canada’ the authors claim to have identified the toxin Cry1Ab in the blood of pregnant women.  Cry1Ab is a protein produced by the bacteria Bacillus thuringiensis (Bt) that is toxic to certain insect pests. Cry1Ab is just one version (event) of this Bt toxin.  Bt toxins have been used extensively by organic farmers and biotechnology has enabled seed companies to develop corn plants that express Cry1Ab proteins giving them a built in defense mechanism against insects susceptible to the toxin, while preserving the biodiversity of friendly insects.  Bt genetics have also been incorporated into cotton. The economic, environmental, safety, and health benefits have made this a very popular  tool used by the majority of family farmers. 

One of the major criticisms of the article was the use of the test used to identify the Cry1Ab protein. In the article the authors state:

‘Cry1Ab protein levels were determined in blood using a commercially available double antibody sandwich(DAS)enzyme-linked immune sorbent assay.’

In previous research, the enzyme-linked immunosorbent assay or ELISA test has been shown to be one of the most unreliable tests for detecting Cry1Ab proteins.  Recall, the CV is relative measure of variation measuring the standard deviation relative to the mean.   It can be used as a metric for risk and reliability (such a consistent yield performance or stock returns).  In the article ‘Comparison and Validation of Methods To Quantify Cry1Ab Toxin from Bacillus thuringiensis for Standardization of Insect Bioassays’ the authors investigate procedures commonly used to identify Cry1Ab.  The authors explain:

“We compared three methods of quantification on three different toxin preparations from independent sources: enzyme-linked immunosorbent assay (ELISA), sodium dodecyl sulfate-polyacrylamide gel electrophoresis and densitometry (SDS-PAGE/densitometry), and the Bradford assay for total protein....The Bradford method resulted in statistically higher estimates than either ELISA or SDSPAGE/ densitometry but also provided the lowest coefficients of variation (CVs) for estimates of the Cry1Ab concentration (from 2.4 to 5.4%). The CV of estimates obtained by ELISA ranged from 12.8 to 26.5%, whereas  the CV of estimates obtained by SDS-PAGE/densitometry ranged from 0.2 to 15.4%....we conclude that standardization of Cry1Ab production and quantification by SDS-PAGE/densitometry may improve data consistency.”

If we look at their reported statistics, we can see for ourselves just how high the CV is on the ELISA test (and therefore how unreliable it is as a method for quantifying Cry1Ab)  compared to other proven methods of quantification.

 
So there you have it. A practical example of an application of a very basic statistic, the coefficient of variation.

References:

Maternal and fetal exposure to pesticides associated to genetically modified foods in Eastern Townships of Quebec, Canada. Reprod Toxicol. 2011 May;31(4):528-33. Epub 2011 Feb 18.
Aris A, Leblanc S.

Comparison and Validation of Methods To Quantify Cry1Ab Toxin from Bacillus thuringiensis for Standardization of Insect Bioassays. Andre´ L. B. Crespo,1 Terence A. Spencer,1 Emily Nekl,2 Marianne Pusztai-Carey,3 William J. Moar,4 and Blair D. Siegfried1* APPLIED AND ENVIRONMENTAL MICROBIOLOGY, Jan. 2008, p. 130–135 Vol. 74, No. 1

A Meta-Analysis of Effects of Bt Cotton and Maize on Nontarget Invertebrates. Michelle Marvier, Chanel McCreedy, James Regetz, Peter Kareiva Science 8 June 2007: Vol. 316. no. 5830, pp. 1475 – 1477

Comparison of Fumonisin Concentrations in Kernels of Transgenic Bt Maize Hybrids and Nontransgenic Hybrids. Munkvold, G.P. et al . Plant Disease 83, 130-138 1999.

Indirect Reduction of Ear Molds and Associated Mycotoxins in Bacillus thuringiensis Corn Under Controlled and Open Field Conditions: Utility and Limitations. Dowd, J. Economic Entomology. 93 1669-1679 2000.

Impact of Bt cotton on pesticide poisoning in smallholder agriculture: A panel data analysis. Shahzad Kouser, Matin Qaim. Ecological Economics
Volume 70, Issue 11, 15 September 2011, Pages 2105–2113

Communal Benefits of Transgenic Corn. Bruce E. Tabashnik  Science 8 October 2010:Vol. 330. no. 6001, pp. 189 - 190DOI: 10.1126/science.1196864

Tuesday, June 19, 2012

The Coefficient of Variation


Coefficient of Variation: A relative measure of variation measuring the standard deviation relative to the mean.  

Recall the standard deviation is the square root of variance.

CV = (standard deviation / mean) * 100

This metric is useful for comparing variables that have different standard deviations and different means.

For instance, let’s assume we have two corn hybrids that average 150 bushels an acre, and hybrid A’s yield has a standard deviation of 10 bushels per acre, while hybrid B’s yield has a standard deviation of 50 bushels an acre.  If we are about to plant one of these hybrids, we know that hybrid A is going to provide us more certainty about our expected yields in the fall.  The empirical rule tells us that about 95% of the time if we plant hybrid A, our yields will be between 130 and 170 bu/acre. If we plant hybrid B, 95% of the time yields will be between 50 and 250 bu/acre.

So, with two hybrids with the same average yield, we know that the one with the lower standard deviation will perform more consistently. But what if they both have different yields and different standard deviations?

Hybrid A:  mean 185 bu;/acre  std dev: 5 bu/acre
Hybrid B: mean: 200 bu/acre std dev: 30 bu/acre

Choosing the most reliably yielding hybrid in this case is not quite so easy. But the CV helps us make the comparison:

Hybrid A: CV = (5/185) *100 = 2.7%
Hybrid B: CV =  (30/200)*100 = 15%

We can see by the CV that hybrid A, even though it yields on average less than hybrid B,  will deliver more consistent results.
 
In the context of finance, we can think of the CV as a measure of relative dispersion that can be used to compare the risks of assets  that have different mean (expected )returns.
Reference: Principles of Managerial Finance.  11th Edition. Lawrence J. Gitman.

Wednesday, August 17, 2011

Basic Concepts: Defining The Mean and Variance

"E[X] is the center of gravity (or centroid) of the unit mass that is determined by the density function of X. So the mean of X is a measure of where the values of the random variable X are" centered…. The variance of X will be a measure of the spread or dispersion of the density of X
" ---Mood, Alexander McFar1ane, INTRODUCTION TO THE THEORY OF STATISTICS  1963.

“The variance of a distribution provides a measure of the spread or dispersion of the distribution about its mean. A small value of the variance indicates that the probability distribution is tightly concentrated around the mean (μ ); and a large value of the variance typically indicates that the probability distribution has a wide spread about the mean (μ )….the variance of X is the second central moment of X” – DeGroot & Schervish. Probability and Statistics 2002.



Monday, May 23, 2011

Type II Error Bias and the FDA

Let’s assume that FDA researchers are reviewing a drug for possible approval. They are reviewing results and data from clinical trials etc. Of course there are some side effects , but only from a small percentage of the samples. Let’s look at a model of their decision making process in the context of hypothesis testing. Let’s assume:

Ho: X = Xo ‘drug is harmful’
Ha: X = Xa ‘drug is safe’

If the decision makers make the mistake of releasing a harmful drug, the consequences would be easily identifiable, and could be visibly traced to their decision. They would want to avoid this at all cost. In essence, they want to avoid making a type I error.

Type I Error: = releasing a harmful drug

From my previous discussion on hypothesis testing, we know that to decrease the probability of committing a type I error, we choose an alpha level that is lower. In a t-test, if we are really afraid of making a type I error, we might set alpha ( the significance level) at .05, .01, or to go overboard.001 etc.

As previously discussed, setting alpha lower and lower increases beta, or the probability of committing a type II error. Recall, a type II error is accepting a false Ho. In this case that would be equivalent to falsely concluding that the ‘drug is harmful’ when it could actually be released and improve the lives of millions.

Type II error bias = over precaution; setting alpha so low, or setting the standard of proof so high as to almost always reject Ho, and biasing the decision in such a way that the probability of a type two error (beta) is greatly increased.

As a result of type II error bias, many life improving drugs never make it to the market.

Type I and Type II Errors

In an experimental setting where a statistical test of a hypothesis is conducted one may either reject the null hypothesis ‘Ho’, or fail to reject Ho. Let’s assume that the true population parameter we are testing is X. Our null hypothesis may be Ho: X=Xo.


The probability of rejecting the null hypothesis when it is true can be defined by:


Alpha = P (reject Ho X = Xo)


Rejecting the null hypothesis when it is actually true is referred to in statistics as a type I error. Therefore alpha is the probability of ‘committing’ a type 1 error. In a basic statistics, when you conduct a basic t- test ( i.e. reject Ho if t > t-critical) at the 5% level of significance, you are establishing a 5% chance of committing a type one error.


Of course, you may fail to reject Ho, or loosely speaking ‘accept’ Ho.


The probability of ‘accepting’ Ho when it is false can be defined by:


Beta = P (accept Ho X = Xa)


Speaking heuristically, accepting a false Ho is referred to as a type II error. Beta is therefore the probability of committing a type 2 error.


It turns out, that as you increase the significance level of a test ( by making alpha lower and decreasing the probability of a type I error), the probability of a type II error (beta) increases. This is the theoretical basis for ‘type II error bias.’

Saturday, March 19, 2011

Applying the Standard Normal Distribution (Z) to Real Data

In a previous post, I introduced the concept of the standard normal variable Z and how it relates to probability. Here I will demonstrate how to translate actual data to Z and make probability interpretations (under the assumption of normality).

Example 1: Assume we are interested in tire mileage, and we know that tires have a population mean mileage μ = 36,500 σ = 5000 x ~ normal.


What is the probability that a tire (x) will last more than 40,000 miles?

To solve this problem, we translate our real word data to data that is distributed standard normal, or we convert x = 40,000 to some value of z and base our probability on z.

Step 1: Calculate Z 

z =  (x-μ)/σ = (40,000-36,500)/5000 = .70

Step 2: Find the corresponding value from the z-table

For z =.70 the value from the table is .7580

Step 3: Interpret and Calculate Probability

The probability of a value > 40,000 is the same as finding a z > .70. Based on the rules we defined for positive z-values, we know that the probability of a value greater than .70 = 1-.7580 = .2420.  This implies that the probability that Z will exceed .70 and that our original data (x =mileage) will exceed 40,000 is 24.2 %.


Example 2: Assume we are interested in tire mileage, and we know that tires have a population mean mileage μ = 36,500 σ = 5000  x ~ normal.


What is the probability that a tire (x) will last less than 30,100 miles?

 To solve this problem, we translate our real word data to data that is distributed standard normal, or we convert x = 40,000 to some value of z and base our probability on z.

Step 1: Calculate Z


z =  (x-μ)/σ = (30,100-36,500)/5000 = -1.28

Step 2: Find the corresponding value from the z-table 

For z = -1.28 the value from the table is .8999 ~ .90

Step 3: Interpret and Calculate Probability 

The probability of a value of z < -1.28 and the probability that x (mileage) will be less than 30,100 is 1-.90 = .10 or 10%.




Example 3: Assume we are interested in tire mileage, and we know that tires have a population mean mileage μ = 36,500 σ = 5000 x ~ normal.


What is the probability that a tire (x) will last less than 50,000 miles?


Step 1: Calculate Z



z =  (x-μ)/σ = (50,000-36,500)/5000 =  2.7


Step 2: Find the corresponding value from the z-table 


For z = 2.7 the value from the table is .9965 


Step 3: Interpret and Calculate Probability 

The probability of z  being less than 2.7 or x (tire mileage) being less than 50,000 is 99.65%. ( remember, for a positive Z, the probability of getting a lower value of Z is obtained from taking the corresponding value directly from the standard normal table)


Example 4: Assume we are interested in tire mileage, and we know that tires have a population mean mileage μ = 36,500 σ = 5000 x ~ normal.


What is the probability that a tire (x) will last between 30,000 and 40,000 miles?

Step 1:  Calculate P( x < 30,000)  

z =  (x-μ)/σ = (30,000-36,500)/5000 = -1.3

Based on the value from the z table, and a negative z value, the probability of getting a lower x (or the probability of mileage < 30,000) = 1-.9032 = .0968 ~ 9.7%

Step 2: Calculate P( x > 40,000)

z =  (x-μ)/σ = (40,000-36,500)/5000 =.70

From example 1 we know this turns out to be 24.2%.

Step 3:  Calculate P(30,000 <= x <= 40,000)

This is the same as finding the probability of a z between -1.3 and .70. We get this by subtracting the results from step 1 and 2:

1 - .097 - .242 = .661 ~ 66.1 % or
100 - 9.7% - 24.2%= 66.1%

Using the Standard Normal Table to Calculate Probability

You may be familiar with the standard normal distribution and how it can be used to make probability interpretations by looking up values from the standard normal table  that correspond to specific values of z. These values from the table represent the area under the standard normal curve, and have a probability interpretation.

There are 3 instances in which you may be looking at z-values. 1) when z is positive, 2) when z is negative, and 3) when you will be looking at ranges between 2 values of z (which could be positive or negative) Below I outline these instances and the rules you will need to know to correctly use the z-table to calculate probabilities.

z = (x-μ)/σ

1) If z is positive:

a) the probability of observing a lower value of z is the area to the left of z and is the value taken directly from the standard normal table.
b) The probability of a larger value of z is the area to the right of z and = 1-(value from the table)

2) If Z is negative:

a) The probability of a lower z is the area to the left of z and  = 1-(value from table)
b) The probability of a higher value of z is the area to the right of z and = the value directly from the table.


3) The probability of getting a value between -z and z is the area between -z and z. You get this by:

a) calculate the area < - Z= A
b) calculate the area > Z =B
c) calculate 1 - A -B

Wednesday, February 9, 2011

Chebyschev, the Empricial Rule, and Few other Basic Concepts



Chebyshev's theorem can be used to make statements about the proportion of data values that must be within a specified number of standard deviations of the mean, regardless of the distribution.
The empirical rule can be used to determine the percentage of data values that must be within one, two, and three standard deviations of the mean for data having a bell-shaped distribution.

Correlation coefficient: measures the strength of the linear relationship between x and y. 

POPULATION PARAMETERS VS SAMPLE STATISTICS
When we take sample data and calculate a mean, we are calculating a sample statistic. The sample mean is used to ‘estimate’ the actual mean of the population we are sampling from. The population mean is referred to as a population parameter. Sample statistics are used to estimate corresponding population parameters. For this reason, sample statistics are often referred to as ‘estimators.’ In statistics we often use different symbols to represent sample statistics vs. sample parameters.
 
For those interested, below is the R code used in tonight's handout.
 


# *------------------------------------------------------------------
# | PROGRAM NAME: EX_CORRELATION_R
# | DATE: 2/8/11
# | CREATED BY: MATT BOGARD 
# | PROJECT FILE:              
# *----------------------------------------------------------------
# | PURPOSE: DEMONSTRATION OF CORRELATION USING R              
# |
# *------------------------------------------------------------------
# | COMMENTS:               
# |
# |  1:  
# |  2: 
# |  3: 
# |*------------------------------------------------------------------
# | DATA USED:               
# |
# |
# |*------------------------------------------------------------------
# | CONTENTS:               
# |
# |  PART 1: positive linear relationship 
# |  PART 2: negative linear relationship
# |  PART 3: no linear relationship
# *-----------------------------------------------------------------
# | UPDATES:               
# |
# |
# *------------------------------------------------------------------
 
 
 
# *------------------------------------------------------------------
# |                
# |PART 1: positive linear relationship
# |  
# |  
# *-----------------------------------------------------------------
 
 
 
x <- c(1,2,3,4,5,6,7,8,9,10)
 
y <- c(2,3,4,4,5,6,9,8,8,10)
 
# plot x and y
 
plot(x,y)  
title("positive linear relationship") 
 
# fit a linear regression line to the data (a topic for later in the semester)
 
reg1 <- lm(y~x) 
print(reg1) # output
abline(reg1) # plot line
title("positive linear relationship")
 
cov(x,y) # covariance between x and y
 
sd(x) # standard deviation of x
sd(y) # standard deviaiton of y
 
cor(x,y) # correlation coefficient for x and y
 
 
# *------------------------------------------------------------------
# |                
# |PART 2: negative linear relationship
# |  
# |  
# *-----------------------------------------------------------------
 
 
 
# let's keep the same x as above, but look at new data for y: 
 
y2 <- c(9,10,8,7,5,4,6,4,2,1) # read in data for y2
 
# plot x and y2
 
plot(x,y2)
 
# fit line to x and y2 
 
reg2 <- lm(x~ y2)
summary(reg2) 
abline(reg2)
title("negative linear relationship")
 
cov(x,y2) # covariance between x and y2
sd(x) # standard deviation of x
sd(y2) # standard deviation of y2
cor(x,y2) # correlation between x and y
 
 
# *------------------------------------------------------------------
# |                
# |PART 3: no linear relationship
# |  
# |  
# *-----------------------------------------------------------------
 
 
y3 <- c(5,7,10,5,1,8,7,4,5,9) # read in y3 data
 
plot(x,y3) # plot x and y3 data
 
# fit line to x and y3
 
 
reg2 <- lm(x~ y3)
summary(reg2) 
abline(reg2)
title("no linear relationship")
 
cov(x,y3) # covariance between x and y3
sd(x) # standard deviation of x
sd(y3) # standard deviation of y3
cor(x,y3) # correlation between x and y3
Created by Pretty R at inside-R.org

Wednesday, February 2, 2011

Commands for Calculating Mean, Variance, & Standard Deviation

Commands in Excel

Statistic: MEAN
Command: =AVERAGE(left click and select data)

Statistic: VARIANCE
Command: =VAR(left click and select data)

Statistic: STANDARD DEVIATION
Command: STDEV(left click and select data)


Commands in R

garst <- c(148, 150, 152, 146, 154)    # enter values for 'garst'

print(garst)     # see if the data is there

var(garst)      # compute variance

sd(garst)      # compute standard deviation

plot(garst)   # plot

hist(garst) # histogram or distribution


Tuesday, January 25, 2011

Data Distributions: The Gaussian Copula & Fat Tails




For a  basic explanation of mortgage backed securities &  toxic assets as they relate to the credit crisis see:

The Credit Crisis Visualized Part 1

The Credit Crisis Visualized Part 2

 From: In defense of the Gaussian copula, The Economist

"The Gaussian copula provided a convenient way to describe a relationship that held under particular conditions. But it was fed data that reflected a period when housing prices were not correlated to the extent that they turned out to be when the housing bubble popped."

Decisions about risk, leverage, and asset prices would very likely become more correlated in an environment of centrally planned interest rates than under 'normal' conditions.

See also: Models and Agents-Hit by a Fat Tail

"Economists talk about fat tails when they want to refer to the probability that extreme events such as the above occur in a much higher frequency than you would regard as "normal."

"So why tails? And why fat?!"
read more to find out.

Some examples of data generated from various copula functions and assumptions (using R) :



Tuesday, May 4, 2010

Inference, Hypothesis Testing, & Regression


Inference: the exercise of providing information about the population based on information contained in the sample

Type 1 Error: rejecting a true null hypothesis
- the probabability of rejecting a true null hypothesis ( or the probability of a type 1 error) is equal to alpha or 'α' This is also the same as the level of significance in a hypothesis test.

Type II Error: failing to reject, or loosely speaking, 'accepting' the null hypthesis when it is false
-the probability of a type II error =beta or β

Power: the probability of rejecting the null hypothesis when it is false. (1-β)

Regression Model: describes how the dependent variable (y) is related to the independent variable (x), also known as the least squares line.

y = b0 + b1*x

This is derived by the process of least squares, which gives the value for the slope (b1) and the intercept (bo) that minimizes the sum of the squared deviations between the observed values of the dependent variable (Y) and the estimated values of the dependent variable (Y*).

Co-efficient of determination(R-squared) = SSR/SST -> gives the proportion of total variation explained by the regression model. Larger values indicate that the sample data are closer to the least squares line, or a stronger linear relationship exists.

Thursday, April 29, 2010

Regression Demo

Tuesday, April 6, 2010

Statistics Definitions

SAMPLE POINT- each individual outcome of an experiment
SAMPLE SPACE-the collection of all possible sample points in an experiment
EVENT-a collection of sample points
PRIOR PROBABILITY-initial estimate of the probability of an event
RANDOM VARIABLE-a numerical description of the outcome of an experiment
EXPECTED VALUE OF A RANDOM VARIABLE-a measure of the average value of a random variable
VARIANCE OF A RANDOM VARIABLE-a measure of the dispersion of a random variable
PARAMETER-a numerical measure from a population
STATISTIC- a numerical measure from a sample
SAMPLING DISTRIBUTION-a probability distribution for all possible values of a statistic
PROPERTIES OF ESTIMATORS:
Unbiased: Property of a statistic where the expected value of a statistic/estimator = the parameter being estimated
Consistency: as n increases the probability that the value of a statistic/estimator gets closer to the parameter being estimated increases
CENTRAL LIMIT THEOREM-the sampling distribution of the sample mean can be approximated by a normal probability distribution as the sample size becomes large. This holds even if the population from which the sample mean comes from is not normal, or unknown.
ASYMPTOTIC DISTRIBUTION- an estimator’s sampling distribution in large samples
ASYMPTOTIC PROPERTIES OF AN ESTIMATOR- properties of an estimator in large samples

Sunday, March 21, 2010

Video: Z-Scores

Sunday, January 31, 2010

Animal Cruelty and Statistical Reasoning

In a recent article, animal rights activists (Mercy for Animals-MFA) went undercover and made some observations about animal abuse on dairy farms. See-
Governor Paterson, Shut This Dairy Down

The author of the above article states:

"But the grisly footage that every farm randomly chosen for investigation--MFA has investigated 11--seems to yield, indicates the violence is not isolated, not coincidental, but agribusiness-as-usual."

Where the statement above could get carried away, is if someone tried to apply it not only to the population of dairy farmers in that state or region, but to the industry as a whole. It's not clear how broadly they are using the term 'agribusiness as usual' but let's say a reader of the article wanted to apply it to the entire dairy industry.

This is exactly why economists and scientists employ statistical methods. Anyone can make outrageous claims about a number of policies, but are these claims really consistent with evidence? How do we determine if some claims are more valid than others?

Statistical inference is the process by which we take a sample and then try to make statements about the population based on what we observe from the sample. If we take a sample (like a sample of dairy farms) and make observations, the fact that our sample was 'random' doesn't necessarily make our conclusions about the population it came from valid.

Before we can say anything about the population, we need to know 'how rare is this sample?' We need to know something about our 'sampling distribution' to make these claims.

According to the USDA, in 2006 there were 75,000 dairy operations in the U.S. According to the activists claims, they 'randomly' sampled 11 dairies and found abuse on all of them. That represents just .0146% of all dairies. If we wanted to investigate the proportion of dairy farms that were abusing animals, if we wanted to be 90% confident in our estimate ( that is construct a 90% confidence interval) and we wanted the estimate (within the confidence interval)to be within a margin of error of .05, then the sample size required to estimate this proportion can be given by the following formula:

n = (z/2E)^2 where

z = value from the standard normal distribution associated with a 90% confidence interval

E = the margin of error

The sample size we would need is: (1.645/2*.05)^2 = (16.45)^2 = 270.65 ~271 farms!

To do this we have to make some assumptions:

Since we don't know the actual proportion of dairy farms that abuse animals, the most objective estimate may be 50%. The formula above is derived based on that assumption. (if we assumed 90% then it turns out based on the math (not shown) that the sample size would have to be the same as if we assumed that only 10% of farms abused their animals, which gives a sample size of about 98 or way more than 11). This also assumes normally distributed data. But to calculate anything, we would have to depend still on someone's subjective opinion of whether a farm was engaging in abuse or not.

I'm sure the article that I'm referring to above was never intended to be scientific, but the author should have chosen their words more carefully. What they have is allegedly a 'random' observation and nothing more. They have no 'empirical' evidence to infer from their 'random' samples that these abuses are 'agribusiness-as-usual' for the whole population of dairy farmers.

While MFA may have evidence sufficient for taking action against these individual dairies, the question becomes how high should the burden of proof be to support an increase in government oversight of the industry as a whole? (which seems to be the goal of many activist organizations)This kind of analysis involves consideration of the tradeoffs involved. This may depend partly on subjective views. We can use statistics to validate claims made on both sides of the debate, but statistical tests have no 'power' in weighing one person's preferences over another. Economics has no way to make interpersonal comparisons of utility.

Note: The University of Iowa has a great number of statistical calculators for doing these sorts of calculations. The sample size option can be found here. In the box, just select 'CI for one proportion' Deselect finite population ( since the population of dairies is quite large at 75,000)then select your level of confidence and margin of error.

References:

Profits, Costs, and the Changing Structure of Dairy Farming / ERR-47
Economic Research Service/USDA Link

"Governor Patterson Shut Down This Dairy", Jan 27,2010. OpEdNews.com

R Code for Basic Histograms

#####################################################
## THIS IS A LITTLE MORE STRAIGHT FORWARD THAN THE ##
## THE OTHER EXAMPLE- USING A SIMPLE DATA SET READ ##
## DIRECTLY FROM R VS. A FILE OR WEB ##
#####################################################

#----------------------------------------#
# SCHAUM'S P. 49 3.5 SALARY DATA
# BASIC HISTOGRAM
#----------------------------------------#


# LOAD SALARY DATA INTO VARIABLE SALARY

salary<-c(240,240,240,240,240,240,240,240,255,255,265,265,280,280,290,300,305,325,330,340) print(salary) # SEE IF IT IS CORRECT library(lattice) # LOAD THE REQUIRED lattice PACKAGE FOR GRAPHICS hist(salary) # PRODUCE THE HISTOGRAM # SEE OUTPUT BELOW

#---------------------------------------
# HISTOGRAM OPTIONS FOR COLOR ETC.
#----------------------------------------

hist(salary, breaks=6, col="blue") # MANIPULATE THE # OF BREAKS AND COLOR

# CHANGE TITLE AND X-AXIS LABEL

hist(salary, breaks=6, col="blue", xlab ="Weekly Salary", main ="Distribution of Salary")

#-------------------------------------------------
# FIT A SMOOTH CURVE TO THE DATA- KERNAL DENSITY
#-------------------------------------------------

d <- density(salary) # CREATE A DENSITY CURVE TO FIT THE DATA

plot(d) # PLOT THE CURVE

plot(d, main="Kernal Density Distribution of Salary") # ADD TITLE

polygon(d, col="yellow", border="blue") # ADD COLOR

Tuesday, January 5, 2010

EXAMPLE CODE

SAS CODE

Histogram


R CODE

NOTE: the code in these posts looks sloppy, but if you cut and paste into notepad or the R scripting window it lines up nice

Histogram

Histogram using Schaum's Homework Data (more straightforward than example above)

Kernal Density Plots -

R Code Mean, Variance, Standard Deviation

Selected Practice Problems for Test 1

Regression using R


Plotting Normal Curves and Calculating Probability

Monday, January 4, 2010

Concepts from Mathematical Statistics

A. Concepts from Mathematical Statistics

Probability Density Functions

Random Variables

A random variable takes on values that have specific probabilities of occurring
( Nicholson,2002). An example of a random variable would be the number of car accidents per year among sixteen year olds.

If we know how random variables are distributed in a population, then we may have an idea of how rare an observation may be.

Example: How often sixteen year olds are involved in auto accidents in a year’s time.

This information is then useful for making inferences ( or drawing conclusions) about the population of random variables from sample data.

Example: We could look at a sample of data consisting of 1,000 sixteen year olds in the Midwest and make inferences or draw conclusions about the population consisting of all sixteen year olds in the Midwest.

In summary, it is important to be able to specify how a random variable is distributed. It enables us to gauge how rare an observation ( or sample) is and then gives us ground to make predictions, or inferences about the population.

Random variables can be discrete, that is observed in whole units as in counting numbers 1,2,3,4 etc. Random variables may also be continuous. In this case random variables can take on an infinite number of values. An example would be crop yields. Yields can be measured in bushels down to a fraction or decimal.

The distributions of discrete random variables can be presented in tabular form or with histograms. Probability is represented by the area of a ‘rectangle’ in a histogram
( Billingsly, 1993).

Distributions for continuous random variables cannot be represented in tabular format due to their characteristic of taking on an infinite number of values. They are better represented by a smooth curve defined by a function ( Billingsly, 1993). This function is referred to as a probability density function or p.d.f.

The p.d.f. gives the probability that a random variable ‘X’ takes on values in a narrow interval ‘dx’ ( Nicholson, 2002). This probability is equivalent to the area under the p.d.f. curve. This area can be described by the cumulative density function c.d.f. The c.d.f. gives the value of an integral involving the p.d.f.

Let f(x) be a p.d.f.

P( a <= X <= b ) = a b f(x) dx = F ( x )

This can be interpreted to mean that the probability that X is between the values of ‘a’ and ‘b’ is given by the integral of the p.d.f from ‘a’ to ‘b.’ This value can be given by the c.d.f. which is F ( x ). Those familiar with calculus know that F(x) is the anti-derivative of f(x).


Common p.d.f’s

Most students are familiar with using tables in the back of textbooks for normal, chi-square, t, and F distributions. These tables are generated by the p.d.f’s for these particular distributions. For example, if we make the assumption that the random variable X is normally distributed then its p.d.f. is specified as follows:

f(x)= 1 / (2 ) -1/2 e –1/2 (x - )2/

where X~ N (  )

In the beginning of this section I stated that it was important to be able to specify how a random variable is distributed. Through experience, statisticians have found that they are justified in modeling many random variables with these p.d.f.’s. Therefore, in many cases one can be justified in using one of these p.d.f.’s to determine how rare a sample observation is, and then to make inferences about the population. This is in fact what takes place when you look up values from the tables in the back of statistics textbooks.

Mathematical Expectation

The expected value E(X) for a discrete random variable can be defined as ‘the sum of products of each observation Xi and the probability of observing that particular value of Xi ( Billingsly, 1993). The expected value of an observation is the mean of a distribution of observations. It can be thought of conceptually as an average or weighted mean.

Example: given the p.d.f. for the discrete random variable X in tabular format.

Xi : 1 2 3

P(x =xi) .25 .50 .25


E (X) = Xi P(x =xi) = 1 * .25 + 2*.50 + 3*.25 = 2.0

Conceptually, the expected value of a random variable can be viewed as a point of balance or center of gravity. In reality, the actual value of an observed random variable will likely be larger or smaller than the ‘expected value.’ However due to the nature of the expected value, actual observations will have a distribution that is balanced around the expected value. For this reason the expected value can be viewed as the balancing point for a distribution, or the center of gravity ( Billingsly, 1993).

That is to say that population values cluster or gravitate around the expected value. Most of the population values are expected to be found within a small interval ( i.e. measured in standard deviations) about the population’s expected value or mean. Hence the expected values gives a hint about how rare a sample may be. Values lying near a population mean should not be rare but quite common.

The expected value for a continuous random variable must be calculated using integral calculus.

E(X) = ∫ x f(x) dx

If x is an observed value and f(x) is the p.d.f. for x, then the product x f(x) dx is the continuous version of the discrete case Xi P(x =xi) and integration is the continuous version of summation.

Variance

As I mentioned previously, actual observations will often depart from the expected value. Variance quantifies the degree to which observations in a distribution depart from an expected value. It is a mean of squared deviations between an observation and an expected value/mean (Billingsly, 1993).

In the discrete case we have the following mathematical description:

X -  )2 = 2

In the continuous case:

∫ (X -  )2 f(x) dx = E( X -  ) = 2


Given the mean (), or expected value of a random variable, one knows the value a random variable is likely to assume on average. The variance (2) indicates how close these observations are likely to be to the mean or expected value on average.




Sample Estimates

Given knowledge of the population mean and variance, one can characterize the population distribution for a random variable. As I mentioned at the beginning of the previous section, it is important to be able to specify how a random variable is distributed. It enables us to gauge how rare an observation ( or sample) is and then gives us grounds to make predictions, or inferences about the population.

It is not always the case that we have access to all of the data from a population necessary for determining population parameters like the mean and variance. In this case we must estimate these parameters using sample data to compute estimators or statistics.

Estimators or statistics are mathematical functions of sample data. The approach most students are familiar with in computing sample statistics is to compute the sample mean ( Xbar) and sample variance (s2) from sample data to estimate the population mean () and variance (2). This is referred to as the analogy principle, which involves using a corresponding sample feature to estimate a population feature (Bollinger,2002).


Properties of Estimators

A question seldom answered in many undergraduate statistics or research methods courses is how do we justify computing a sample mean from a small sample of observations, and then use it to make inferences about the population?

I have discussed the fact that most of the values in a population can be expected to be found within a small interval about the population mean or expected value. The question remains to be can we expect that most of the values of a population observation be found within a small interval of a sample mean? This must be true if we are to make inferences about the population using sample data and estimators like the sample mean.

Just like random variables, estimators or statistics have distributions. The distributions of estimators/statistics are referred to as sampling distributions. If the sampling distribution of a statistic is similar to that of the population then it may be useful to use that statistic to make inferences about the population. In that case sample means and variances may be good estimators for population means and variances.

Fortunately statisticians have developed criteria for evaluating ‘estimators’ like the sample mean and variance. There are four properties that characterize good estimators.




Unbiased Estimators

If ^ is an estimator for the population parameter , and if
E (^ then ^ is an unbiased estimator of the population parameter . This implies that the distribution of the sample statistic/estimator is centered around the population parameter ( Bollinger, 2002).

Consistency

^ is a consistant estimator of  if

lim as n infinity : Pr[ | ^ -  | < c ] = 1

This implies that as you add more and more data, the probability of getting closer and closer to  gets large or variance of ^ approaches zero as n approaches infinity. It can be said that ^ p  or converges in probability to  (Bollinger, 2002).

Efficiency

This is based on the variance of the sample statistic/estimator. Given the estimator ^, and an alternative estimator ~, ^ is more efficient given that

Variance(^) < Variance(

Mean Squared Error is a method of quantifying efficiency.

MSE = E [ [ - ^]2] = V(^) + E[ (E(^) - )]2 = variance + bias squared.

It can then be concluded that for an unbiased estimator the measure of efficiency reduces to the variance (Bollinger, 2002).

Robustness

Robustness is determined by how other properties ( i.e. unbiasedness,consistency,efficiency,) are affected by assumptions made
(Bollinger, 2002).

When sample statistics or estimators exhibit these four properties, statisticians feel that they can rely on these computations to estimate population parameters. It can be shown mathematically that the formulas used in computing the sample mean and variance meet these criteria.






Confidence Intervals

Confidence intervals are based on the sampling distributions of a sample statistic/estimator. Confidence intervals based on these distributions tell us what values an estimator ( ex: the sample mean) is likely to take, and how likely or rare the value is ( DeGroot, 2002). Confidence intervals are the basis for hypothesis testing.

Theoretical Confidence Interval

If we assume that our sample data is distributed normally, Xi ~ N (2) then it can be shown that the statistic

Z = ( Xbar - )2 / (2 / n)1/2 ~ N( 0, 1) Standard Normal Distribution ( Billingsly, 1993).

Given the probabilities represented by the standard normal distribution it can be shown as a matter of algebra that

Pr ( -1.96 <= Z <= 1.96) = .95

The value 1.96 is the qauntile of the standard normal distribution such that there is only a .025 or 2.5% chance that we will find a Z value greater than 1.96. Conversely there is only a 2.5% chance of finding a computed Z value less than –1.96
( Steele, 1997).

As a matter of algebra it can be shown that


Pr(Xbar - 1.96 (2 / n)1/2 <= <= Xbar + 1.96 (2 / n)1/2 ) = .95

( Goldberger, 1991).

This implies that 95% of the time we can be confident that the population mean will be 1.96 standard deviations from the sample mean. The interval above then represents a 95% confidence interval ( Bollinger, 2002).


The Central Limit Theorem –Asymptotic Results

The above confidence interval is referred to as a Theoretical confidence interval. It is theoretical because is based on knowledge of the population distribution being normal. According to the central limit theorem:

Given random sampling, E(X) = , V(X) = 2 , the Z- statistic Z = ( Xbar - )2 / (2 / n1/2) converges in distribution to N( 0,1). That is


( Xbar - )2 / (2 / n1/2) ~A N(0,1)

It can then be stated that the statistic is asymptotically distributed standard normal.
Asymptotic properties are characteristics that hold as the sample size becomes large or approaches infinity. The CLT holds regardless of how that sample data is distributed, hence there are no assumptions about normality necessary (DeGroot 2002).

Student’s t Distribution-Exact Results

A limitation of the central limit theorem is that it requires knowledge of the population variance. In many cases we use s2 to estimate 2 . Gosset, a brewer for Guiness in the 1800’s was interested in normally distributed data with small sample sizes. He found that using Z with s2 to estimate 2 did not work well with small samples. He wanted a statistic that relied on exact results vs. large sample asymptotics ( Steele, 1997).

Working under the name Student, he developed the t distribution, where

t = ( Xbar - )2 / (s2 / n)1/2 ~ t(n-1)


The t distribution is the ratio of a normally distributed variable and chi-square distributed variable ( DeGroot, 2002). It is important to note that the central limit theorem does not apply because we are using s2 instead of 2. Here we can rely on using the t-table for constructing confidence intervals and rely on exact results vs. the approximate or asymptotic results of the
CLT ( DeGroot, 2002).

More Asymptotics- Extending the CLT

Sometimes we don’t know the distribution of the data we are working with, or don’t feel comfortable making assumptions of normality. Usually we have to estimate 2 with s2 . In this case we can’t rely on the asymptotic results of the CLT or the exact results of the t-distribution.

In this case there are some powerful theorems regarding asymptotic properties of sample statistics known as the Slutsky Theorems.