Showing posts with label Statistics. Show all posts
Showing posts with label Statistics. Show all posts

Saturday, 12 November 2011

Binomial Approximation


A Bin(n,p) distribution is the sum of n Bernoulli variables, so the Central Limit Theorem says that it will approximate to a Normal distribution if n is large enough. n could be a few dozen, less for a symmetrical distribution and more for a very non-symmetrical distribution.


For the Poisson distribution E[X] = var[X] = μ. A Bin(n,p) distribution has E[X] = np and var[X] = npq. So for a Binomial distribution to approximate to a Poisson distribution np ≈ npq, i.e. q ≈ 1. For example, if n = 1000 and p = 0.01, then E[X] = 10 and var[X] = 9.9, and a Poi(10) distribution could be considered as an approximation, depending on any other relevant criteria.


The Bin(n,p) distribution has an upper limit of n, while neither the Normal nor Poisson distributions have an upper limit.

Memoryless Property


Basically it means the distribution doesn't know where it is on the curve. Formally P(X > a | X > a-b) = P(X > b)
The only dists in CT3 with this property are exponential and geometric.
eg. if X~Exp then P(X>4 | X>3) = P(X>1).

Hypothesis Testing


Suppose our null hypothesis is that the mean age of students = 22 and our alternative hypothesis is that the mean age is not equal to 22.

Suppose we take a sample and find that the sample mean age is 26.

We can deal with this in two main ways:

(1) What is the probability of getting this sample result if Ho is true? This is called the probability method and the probability is called the p-value.

(2) What value of the sample mean would convince me that Ho is true/false? This is called the critical value method.

Using (1), we can calculate the probabilty of getting a sample mean as extreme as this on both sides (ie >26 or <18) if the population mean is actually 22. Suppose this probability turns out to be large, eg 34%. We would say that getting a sample mean as extreme as this if the population mean is actually 22 is not very unusual at all so therefore we would say that we do not have sufficient evidence to reject Ho. However, if the probability of getting a sample mean as extreme as this if the population mean is actually 22 is very low, eg 3%, we would say that this is very unusual and therefore we would reject Ho.

We usually use 5% as a cut-off point, ie we usually run a 5% Type I error - so there is a 5% chance of rejecting Ho when it is in fact true. This is called the significance level of the test.

Using (2) we can determine the critical values of the sample mean, C1 and C2 (or the test statistic, eg Z1 and Z2) such that the probability of a result lower than C1 or higher than C2 is 5%. Then if our sample mean (or the test statistic) is in this critical range, we will reject Ho at the 5% level.

Negative Binomial Distribution


Differences between Type I and Type II Negative Binomial Distribution and how to identify which is it?
Type 1 counts the number of trials up to and including the kth success.

Type 2 counts the failures before the kth success.

So suppose we are counting the number of goals we score (success) for the penalty kicks we make (trials).

Suppose we score our 3rd goal on our 10th penalty kick.

Under Type 1 X=10 as it is the 10th kick when we got our 3rd goal.

Under Type 2 X=7 as we had 7 failures before we got our 3rd goal.

In CT3 our default choice will be a type 1 (for CT6 it will be type 2).

Source -Nick Campbell