Foreword¶
I will now comprehensively cover the essential probability theory that you need to be equipped with to be able to handle the rest of the material. I will start with convergence in probability.
Modes of Convergence¶
Let us fix our triple . Random variables are measurable functions from this space to . The key intuition to understanding these definitions is to try coming up with them. We want to have some convergence definitions of random variables. But we know real analysis, and we want to involve the sample space in the definition. Hence it is natural to define convergence of random variables in terms of the measure of events which describe their convergence.
Almost Sure convergence: An event happens almost surely if the measure of the set in which the event does not happen is 0. Let us use this for convergence.
This convergence is in relation to pointwise convergence event.
Convergence in Measure: A sequence converges to if for every . Think about this definition a bit. We are calculating the measure of an event, and we need to have some convergence property here. Hence the necessity of such a “contrived” definition.
I have to be clear about the notation here. By , I mean , since that is where our measure is defined. I will not explain this later and continue on with the abuse of notation. I would also like to point out that this makes sense because our triple is fixed. More generally, we can have the random variables codomain be different spaces, but our preimage must be the same measure space for each random variable.
This convergence is in relation to the function distance. (To talk about function convergence, one can look at pointwise convergence, or one can think about the function difference and consider some norm there. Since we are dealing with real numbers here, we took the standard norm on the real numbers).
convergence: in , if .
Framed another way, this is .
This is using the metric and considering the expectation instead. We will always be considering , as that is when it is a metric space.
There is another convergence, called convergence in distribution/convergence in law/weak convergence, which we will cover shortly. This will require a much more detailed conversation.
I will now refer to the convergences by their indices.
Moreover, convergence in measure implies that there is a subsequence that converges almost surely (this is the version of the converse of ), and convergence in implies convergence in if .
Let us state them precisely and prove them (except 4).
Proposition 1: On a probability space, almost surely, then in probability. The converse need not be true.
Proof: Fix . We need to show that . Let
For the coounterexample, consider with lebesgue measure and , and .
Proposition 2: If in measure, then there is a subsequence of which converges to a.e.
For each , . Fix such that the measure is less than . Let . Now we are done by Borel Cantelli Lemma.
Proposition 3: Let . If in , then in measure.
This follows directly from Chebyshev.
where the norm is the norm.
Propostition 4: Let . If in , and , then in ,
Weak Convergence
We will discuss this now.
Convergence in Expectation¶
We discussed three convergence modes, and will discuss one more now. But when does convergence of random variables imply that their expectations will converge as well? Turns out, this only happens in convergence.
One counterexample kills all the three other cases. On , take . For each , is eventually 0, hence it converges to 0 almost surely, hence in probability as well as distribution, yet for every .
(For the last case, convergence implies convergence which is convergence of expectations).
The random variables have the property of their expectations converging with one additional required property: Uniform Integratibility. This result is called Vitali theorem.
Vitali Theorem for convergence in probability (hence for almost surely):
If in probability (and are integrable) then
Note that if we have the same statement for , then we have convergence.
Vitali Theorem for weak convergence:
Let . If are uniformly integrable, then and . The converse holds if the means converge for .
Why did I not just repeat the same theorem as the previous case? Because here, you possibly have different sample spaces as preimages for each .
Let me now state the uniform integrability definition.
Definition: A family of real valued random variables is uniformly integrable if
Notice that is uniform over all . This condition says that the tails uniformly carry neglible mass. Also see that this condition is on the laws alone.
Here is an important theorem (de la Valle Poussin):
Definition: A sequence of nonnegative random variables is said to satisfy the asymptotic expectation condition with index if there exists such that
Weak Convergence¶
Consider a probability space .
Almost sure convergence requires knowing almost all sample points. This notion is called convergence with probability one. Strong Law of Large Numbers and Kolmogorov Three Series (not necessary for now) Theorem are results of this convergence.
Convergence in probability does not require knowledge of the values of the variables at individual sample points. It needs the knowledge of certain probability of events. The Weak Law of Large Numbers is a result of this type of convergence.
Convergence in requires the functions in .
In practice, we do not have minute knowledge about random variables. We only perceive expectations, or sum total of certain characteristics. In practice, we often perceive the moments of a random variable, which are meaninful only when they exist. However, if you take any bounded continuous function on , is always meaningful and can be regarded as a statistic. Here is a nice result which makes this significant as a statistic.
Result: If we have two random variables and , and if for every bounded continuous function , then and have the same distribution.
This leads to the following definition:
in distribution if for every bounded continuous function . This is also called convergence in law, denoted
Remark: The first three convergence modes, the measure in the preimage (sample space) matters. For the convergnce in distribtion, only the pushforwarsd of matter by the random variables. These are the measures on , also called the law. The sample space is irrelevent. That is why you can have convergence in distribution even if the random variables are wildly different, hence it is a weak form of convergence. This distinction is important when it comes to convergence, and is underemphasized generally. One can do this precisely because of the change of measure (also called the law of unconcious statistician, now you know why):
where is the pushforward measure on .
A sequence of probabilities on converges weakly to a probability if for every bounded continunous function This is denoted .
We now prove the result. We first use change of measure on the condition, to get two probability distributions (which are pushforwards) and such that
for every bounded continuous function on .
We need to show for every borel set . The proof is quite simple. Clearly, indicator function of as the candidate of is the best, unfortunately it is not continuous. So what? Just take a continuous sequence of functions converging to the indicator function and then use Dominated Convergence theorem to take limit inside. We are done.
Alternatively, another simple proof: Notice that and are bounded continuous functions. Now invoke the uniqueness of characteristic functions to conclude.
By the way, according to the first proof, we can generalize the result to restrict the set of functions to functions on a compact support.
Exercises¶
This is a good time to take a break from theory. I now suggest you to go through the plethora of examples (or exercises) on this topic. They are in another section for this topic available in the notebook.
Theorems in Weak Convergence¶
We need to develop more theory in weak convergence, since looking at each bounded continuous function seems like quite an impractical job to do. Here is a nice toolkit.
Theorem 1: Let and be random variables. and . Distribution functions of and are and respectively. Following statements are equivalent:
.
.
for every continuity point of .
for a dense set of points in .
The second statement reduces our job by transferring the entire action to the real line. Now we only need to calculate integrals on . The third statement reduces our job further. We need not look at every bounded continuous function, just evaluate distribution at a lot points and show convergence of thos numbers. The final statement reduces it even furhter. We do not need evaluation at uncountably many points, just a nice dense set (such as , or some other set) works.
I will hold off on proving this in this notebook, but the reader can look at other resources which prove this. It is just heavy work in analysis.
Theorem 2: The above four conditions are also equivalent to the folowing.
for all bounded functions which have bounded derivatives.
for all real bounded unifromly continuous .
Theorem 3: The above six conditions are also equivalent to the following:
For the next theorem, recall:
This is the characteristic function of , or its distribution function . It is also the fourier transform of . This is a complex valued function whose domain is .
Theorem 4 (Levy’s Continuity Theorem): The above nine conditions are equivalent to the following:
for each .
uniformly over each bounded interval.
This theorem is pretty peak, you can use it to prove the Central Limit Theorem taught in the introductory courses.
Theorem 5 (Skorokhod Theorem): The above eleven conditions are equivalent to the following:
There is a sequence of random variables and on such that for each , , and a.e wrt .
If the above conditions hold, then a.e wrt , and thus the expectations converge (by DCT). But and , thus we are done with the forward direction. For the converse, one needs to look at the quantile function / inverse distribution function.
Theorem: Weak Convergence is robust. and is continuous, and , then .