\[ \renewcommand{\P}{\mathop{\mathbb{P}}\nolimits} \newcommand{\E}{\mathop{\mathbb{E}}\nolimits} \newcommand{\var}{\mathop{\rm Var}\nolimits} \newcommand{\VaR}{\mathop{\rm VaR}\nolimits} \newcommand{\cte}{\mathop{\rm CTE}\nolimits} \newcommand{\cov}{\mathop{\rm Cov}\nolimits} \newcommand{\limsup}{\mathop{\rm limsup}} \newcommand{\liminf}{\mathop{\rm liminf}} \newcommand{\R}{\mathbb{R}} \newcommand{\Q}{\mathbb{Q}} \newcommand{\Z}{\mathbb{Z}} \newcommand{\N}{\mathbb{N}} \newcommand{\C}{\mathbb{C}} \renewcommand{\d}{\, \mathrm{d}} \newcommand{\dP}{\, \mathrm{d}\mathbb{P}} \newcommand{\eps}{\varepsilon} \renewcommand{\emptyset}{\varnothing} \]
8 Sharing and aggregating risks 2
8.1 Why?
In Chapter 7 we separately introduced the concepts of sharing a risk between two (or more!) parties, as well as the idea of aggregating multiple (independent and identical) risks into one single total aggregate risk taking the form of a compound distribution. This brings us naturally to the following situation: a risk manager facing multiple (independent and identical) risks looking to share all of them. What effect does such a sharing agreement have on the aggregate risk for the risk manager? And what does the total risk look like for the other party in the sharing agreement? It turns out that the aggregate/total risk for each party again takes the form of a compound distribution, and hence can very neatly be analysed using the tools we already developed in Chapter 7!
We can even take this one step further. Imagine a risk manager dealing with multiple (independent and identical) risks from several sources, say source A and B, and that the risks are not exchangeable between these sources. Then we could write down a compound distribution for the aggregare risk from source A, say \(S_A\), and similarly \(S_B\) for source \(B\). The total risk for the risk manager is then given by the sum of these two compound distributions i.e. \(S:=S_A+S_B\). What can we say about this \(S\)? Well it turns out, as we will discuss in Section 8.3, that as long as we are dealing with compound Poisson distributions then there is a beautiful mathematical result telling us that \(S\) is also a compound Poisson distribution — and hence that again brings all the tools from Chapter 7 into the play for analysing and understanding this \(S\)! :).
8.2 Risk sharing within the context of aggregate risks/compound distributions
We pick up the setting from Section 7.3 again: you as risk manager own a possibly large number of potential risks/losses that are identical (enough) and can be considered independent of each other, and each of which may or may not materialise. Denoting by \(N \in \{0,1,\ldots\}\) the (random) number of them that materialise and by \(X_1, \ldots, X_N\) the (random) sizes of the risks/losses that turn out to materialise, the (random) total/aggregate risk/loss \(S\) can (hence) be written as \[S=\sum_{i=1}^N X_i \tag{8.1}\] and it has a compound distribution as defined in Definition 7.1. Recall that a crucial element is that we translate “the risks/losses that are identical (enough) and can be considered independent of each other” into the mathematical assumption that \(X_1, X_2, \ldots\) is a sequence of i.i.d. (independent and identically distributed) random variables.
Now imagine that you are looking for a risk sharing agreement as discussed in Section 7.2. This could be at the level of sharing the total risk/loss \(S\), but also at the level of the sharing each of the individual risks that actually materialise i.e. sharing the losses \(X_1, \ldots, X_N\). The former is less interesting and is mathematically essentially the same as the context in Section 7.2: a risk sharing agreement applied to a single random variable, in this case \(S\).
So let’s focus on the more interesting situation of a risk sharing agreement applying to the individual risks that materialise. Make sure to keep the big picture crystal clear for yourself, as follows (extending Proposition 7.1):
Proposition 8.1 The process of sharing a whole collection of potential risks between two parties, with total risk/loss \(S\) as given by Equation 8.1 works as follows:
- a sharing agreement is put in place i.e. functions \(f_1\) and \(f_2\) are chosen as discussed in Section 7.2, and possibly (premium) payments made between the parties involved;
- over time, as the risks materialise one after the other, the random variables \(X_1, X_2, \ldots, X_N\) each take their value;
- once \(X_i\) takes its value, it is split up in the parts \[Y_i := f_1(X_i) \quad \text{and} \quad Z_i := f_2(X_i)\] to be paid by each party.
So, after the period is over, party 1 has paid in total \[S_1 := \sum_{i=1}^N Y_i \tag{8.2}\] while party 2 has paid in total \[S_2 := \sum_{i=1}^N Z_i \tag{8.3}\] (with the usual understanding that if \(N=0\), then \(S_1=0\) and \(S_2=0\)).
Of course, in total both parties together have paid the total risk \(S\): \[S_1+S_2=\sum_{i=1}^N Y_i+\sum_{i=1}^N Z_i=\sum_{i=1}^N \left( Y_i+Z_i \right)=\sum_{i=1}^N X_i=S,\] where we used that \[Y_i+Z_i=f_1(X_i)+f_2(X_i)=X_i,\] because \(f_1\) and \(f_2\) must satisfy the relationship Equation 7.2.
Obviously, the interesting perspective is not so much looking back at the previous year (say), but looking forward to a coming year (say) in which steps 2 and 3 of the above process still have to take place. So party 1 resp. party 2 is facing the random risk/total loss Equation 8.2 resp. Equation 8.3, likely hoping that it won’t take too high a value!
If you think a little bit about it, then you’ll agree with me that we can immediately write down some powerful conclusions about both \(S_1\) and \(S_2\): given that \(X_1, X_2, \ldots\) are i.i.d. random variables and independent of \(N\), these properties remain in tact if we apply a function to these \(X_i\)’s. Hence both \(S_1\) and \(S_2\) satisfy the criterions for having a compound distribution, with in particular as bonus that we immediately get all the machinery from Section 7.3.2 to throw at analysing them!
Proposition 8.2 Consider a sharing agreement for a collection of risks as laid out in Proposition 8.1, with the standing assumption that \(X_1, X_2, \ldots\) are i.i.d. random variables independent of \(N\), which has a counting distribution (cf. Definition 7.1).
Then both \(S_1\) given by Equation 8.2 and \(S_2\) given by Equation 8.3 have a compound distribution as defined in Definition 7.1, and hence both Proposition 7.2 and Theorem 7.1 can be used to study their properties: probabilities, moments, mgf’s etc.!
And this is not even all the compound distribution fun we can have in this context, indeed there are even more compound distributions lurking around! Consider for simplicity the original setup without sharing, in which the total risk/loss is given by Equation 8.1 — the argument we’re going to discuss here can be applied to any compound distribution, so including \(S_1\) and \(S_2\) from the above Proposition 8.2 for example.
Imagine that the owner of the risks is interested in a question of the type: how many of the risks materialise and result in a loss of more than \(10\)? Or result in a loss of less than \(100\)? Slightly more generally: how many of the risks materialise and take a value in some set \(B \subseteq \R\)?
Of course the first task is to write down some kind of expression for this number, let’s denote it by \(\widehat{N}\). The (neat, I think!) idea is as follows. Recall that \(N\) risks materialise, with losses \(X_1, \ldots, X_N\). Imagine attaching a label to each loss, with the value \(1\) if \(X_i \in B\) indeed holds and value \(0\) if \(X_i \in B\) does not hold. This gives a sequence of \(N\) (random) labels that we can mathematically express in the form \[I_i = \begin{cases} 1 & \text{if } X_i \in B \\ 0 & \text{if } X_i \not\in B \end{cases} \quad \text{for } i=1,\ldots,N \] (note that each of these \(I_i\)’s is nothing but a good old indicator function, cf. Remark A.1!).
Now, the (random) number \(\widehat{N}\) we’re interested in is clearly equal to the (random) number of labels that get the value \(1\), which will result in some value \(0,1,\ldots,N\). Here is the neat observation: in order to count the number of \(1\)’s in a sequence of \(0\)’s and \(1\)’s, we can simply compute the sum of all the elements in the sequence! For instance, the sequence \(0,1,1,0,0,1\) contains 3 \(1\)’s and indeed its sum is also 3. Hence we can express \(\widehat{N}\) as the sum of our random labels: \[\widehat{N}=\sum_{i=1}^N I_i. \tag{8.4}\] Further, since each \(I_i\) is obtained by applying the function \[g(x)=\begin{cases} 1 & \text{if } X_i \in B \\ 0 & \text{if } X_i \not\in B \end{cases}\] to \(X_i\), by the same argument as above Proposition 8.2 the \(I_i\)’s form an i.i.d. sequence of random variables independent of \(N\), and hence the expression Equation 8.4 is (again) a compound distribution!
Proposition 8.3 Let \(S\) have a compound distribution as defined in Definition 7.1 i.e. be of the form \[S=\sum_{i=1}^N X_i.\] Take some set \(B \subseteq \R\) and define the i.i.d. sequence of random variables \(I_1, I_2, \ldots\) as \[I_i = \begin{cases} 1 & \text{if } X_i \in B \\ 0 & \text{if } X_i \not\in B \end{cases} \quad \text{for all } i=1,2,\ldots \] (hence indicator functions, cf. Remark A.1). Further define \[\widehat{N}:=\sum_{i=1}^N I_i\] (with the usual understanding that \(\widehat{N}=0\) if \(N=0\)).
Then \(\widehat{N}\) counts the number of \(X_i\)’s that take their value in \(B\), and it has a compound distribution as well. Hence both Proposition 7.2 and Theorem 7.1 can be used to study the properties of \(\widehat{N}\): its probabilities, moments, mgf’s etc.
I think that some examples to see this in action are definitely called for — let’s first do one that links back to Section 7.3 in particular, and then one following up that digs more into the above.
Example 8.1 Imagine that you are working for an electronics company, managing the “after sales” responsibility for a large collection of new TVs that have recently been sold. The buyers of these TVs are entitled to claim a partial refund if their TV has an issue. Suppose that if such a claim is made, the amount that is refunded is either £100 (in 10% of cases), £500 (in 30% of cases) or £1,000 (in the remaining 60% of cases). Further you model the total number of refunds that you have to do by a \(\text{Poisson}(10)\) distribution.
First convince yourself how this situation naturally fits into the setup of a compound distribution as discussed in Section 7.3: you are dealing with a large number of potential risks/losses each of which may materialise (each buyer in your collection could end up submitting a claim but normally not all will). The number of buyers that actually does is uncertain/random and modelled by \(N \sim \text{Poisson}(10)\).
Further, the uncertain/random sizes of the risks that actually materialise, which we denote by \(X_1, X_2, \ldots\) are identically distributed with the following common distribution: \[X_i = \begin{cases} 100 & \text{with prob } 0.1 \\ 500 & \text{with prob } 0.3 \\ 1000 & \text{with prob } 0.6. \end{cases} \tag{8.5}\]
Your total risk/loss is then given by \[S=\sum_{i=1}^N X_i, \tag{8.6}\] with the usual understanding that \(S=0\) if \(N=0\) (no claims at all, then no loss at all!). And if we look at Definition 7.1, then we have mentioned almost all conditions needed for \(S\) to have a compound distribution — the only one outstanding is that the \(X_i\)’s also need to be independent. Let’s make that assumption as well.
Note that \(S\) now has a compound Poisson distribution (cf. Definition 7.2), with parameter pair \((10,F_X)\), where \(F_X\) is the cdf of the distribution Equation 8.5.
If we compare this situation with the string of examples for compound distributions in Chapter Chapter 7 (cf. Example 7.3, Example 7.4, Example 7.5 and Example 7.6), then a key difference here is that the \(X_i\)’s now have a discrete rather than a continuous distribution. But that doesn’t spoil the fun! :).
Model criticism. A model is of course always (just) that: a model i.e. a mathematical representation of the actual real world situation we are dealing with. And typically a model is not an exact representation, for instance because we have only a limited understanding of the situation at hand. But quite often it is actually a deliberate choice: there can be a lot of value in a relatively simple model that allows for efficient and (almost) exact analysis. The trade-off is then that you know that you’re causing a modelling error, but if the alternative is a more complicated model in which for instance only approximate analysis is possible then you may well assess that the modelling error from the simple model is actually less severe than the error due to the more complicated model allowing for approximate analysis only!
In this case, I think that you could have two points of criticism of this model:
- We gave \(N\), the number of risks that will materialise, a \(\text{Poisson}(10)\) distribution. That means that in our model, \(N\) has range \(\{0,1,\ldots\}\) i.e. can take an arbitrarily large value! Obviously that is not entirely consistent with the real world: the number of TVs that you are responsible for may be large but it is definitely finite and logically the number of risks that will materialise cannot be larger than that number of TVs. However a Poisson distribution is attractive because it has that natural (for the problem at hand) interpretation also discussed in Section 7.3.1 and as it is such a well known distribution, many of its properties are readily available to us. Additionally, the probability that a Poisson distribution takes a value larger than \(n\) decays quite fast as \(n\) gets “very large” which in combination with the boundedness of the \(X_i\)’s gives good reason to believe (short of a rigorous analysis) that the modelling error we’re causing here is not too bad.
- We assumed that the \(X_i\)’s are independent. We have only limited information at hand, but we did warn in the Copulas part of the course against making that assumption all too easy, and saw how severe the consequences can be if you are really wrong at this front (recall Example 3.2 e.g., and we also touched upon this issue in Exercise 6.5). Independence has the benefit that it makes the analysis a lot easier, so for the same reason as under 1, if we trust that we don’t make too much of an error here then it is preferable.
The probability that you suffer no loss at all. Note that you suffer no loss at all if and only if \(N=0\), after all if \(N \geq 1\) then \(S\) in Equation 8.6 is the sum of one or more random variables that all take a strictly positive value and hence is necessarily strictly positive itself. Or, in words: if one or more refund has to be done, then you’ll definitely have some loss, the only way to have no loss at all is if there are no refunds at all. Hence the prob is \[\P(N=0)=e^{-10} \approx 0.00005\] (since \(N \sim \text{Poisson}(10)\), and of course you can use its pmf from Appendix B if needs be).
The expectation and variance of your total risk/loss. Since \(S\) has a compound distribution, we have all the amazing tools from Section 7.3.2 at our disposal. For the expectation and variance, we can appeal to Theorem 7.1: \[\E[S]=\E[N] \E[X_1] \quad \text{and} \quad \var(S)=\var(N) \big( \E[X_1] \big)^2 + \E[N] \var(X_1).\] Since \(N \sim \text{Poisson}(10)\), from Appendix B we immediately get that \(\E[N]=10\) and \(\var(N)=10\). Further we can compute the expectation and variance of \(X_1\) using the standard formulas for the expectation of (a function of) a discrete random variable (cf. Equation A.10 and Equation A.9, plus Equation 8.5 of course): \[\E[X_1]=100 \cdot 0.1 + 500 \cdot 0.3 + 1000 \cdot 0.6=760\] and \[\E[X_1^2]=100^2 \cdot 0.1 + 500^2 \cdot 0.3 + 1000^2 \cdot 0.6=676000,\] so \[\var(X_1)=\E[X_1^2]- \big( \E[X_1] \big)^2=676000-760^2=98400.\] So we end up with \[\E[S]=\E[N] \E[X_1]=10 \cdot 760=7600\] and \[\begin{align*} \var(S) &= \var(N) \big( \E[X_1] \big)^2 + \E[N] \var(X_1) \\ &= 10 \cdot 760^2+10 \cdot 98400 \\ &= 6760000. \end{align*} \]
The moment generating function (mgf) of your total risk/loss. Let’s also take a look at the mgf \(M_S\) of \(S\), because why not! As for the expectation and variance above, Theorem 7.1 gives us a ready-to-roll expression: \[M_S(t)=M_N \big( \log(M_X(t)) \big) \tag{8.7}\] for all \(t \in \R\) for which \(M_X(t)<\infty\), where \(M_N\) resp. \(M_X\) is mgf of \(N\) resp. the common mgf of the \(X_i\)’s.
Recall that we can pick up \(M_N\) from Definition 7.2 (with \(\lambda=10\)): \[M_N(t)=\exp \left( 10 (e^t-1) \right) \quad \text{for all} t \in \R. \tag{8.8}\] For \(M_X\), we’ll have to work out that one by ourselves, from the common distribution of the \(X_i\)’s given in Equation 8.5. Recall that we did such an mgf comp once before already, in Example 7.6, however there the \(X_i\)’s were continuous while here they are discrete. The idea is exactly the same though: fix a \(t \in \R\), write down the def of the mgf from Section A.2, observe that is an expectation and so work this out using the standard formula for the expectation of a function of a discrete random variable (cf. Equation A.9, with the choice \(h(x)=e^{tx}\) and note that the range of \(X_1\) is \(R:=\{100, 500, 1000\}\)): \[\begin{align*} M_X(t) &= \E \left[ e^{t X_1} \right] \\ &= \E \left[ h(X_1) \right] \\ &= \sum_{k \in R} h(k) \P(X_1=k) \\ &= h(100) \cdot 0.1 + h(500) \cdot 0.3 + h(1000) \cdot 0.6 \\ &= 0.1 e^{100t}+0.3 e^{500t}+0.6 e^{1000t}. \end{align*} \] Easy enough no? Now plugging this together with Equation 8.8 back into Equation 8.7 yields \[\begin{align*} M_S(t) &= M_N \big( \log(M_X(t)) \big) \\ &= \exp \left( 10 \left( e^{\log(M_X(t))}-1 \right) \right) \\ &= \exp \left( 10 \left( M_X(t)-1 \right) \right) \\ &= \exp \left( e^{100t}+3 e^{500t}+6 e^{1000t}-10 \right) \end{align*} \] for all \(t \in \R\).
Just for fun: recall that having the mgf of a random variable gives access to all its moments by means of differentiation (cf. Equation A.20). Applying this, since \[\begin{align*} M_S'(t) &= \left( 100 e^{100t}+1500 e^{500t}+6000 e^{1000t}-10 \right) \\ & \qquad \cdot \exp \left( e^{100t}+3 e^{500t}+6 e^{1000t}-10 \right) \end{align*} \] we get from Equation A.20 that \[\E[S]=M_S'(0)=100 +1500 +6000=7600,\] indeed the same value we had already computed above by using the easier expression provided by Theorem 7.1!
The looooong right tail of \(S\) (fyi only). No \(S\) is not a dog and Kevin doesn’t want to play with \(S\). At all. What are you going to do with a dog with a weird right tail? What we mean here is the following. We saw above that our expected loss is £7,600. However we also see that \(S\) has an impressively large variance — which, recall from your earlier probability courses, is a measure for how “spread out” the values that \(S\) can take are around its expectation (roughly speaking). This is an indication of the type of behaviour you typically see in real world applications of aggregating risks: it tends to have a “long right tail” i.e. compared with a distribution that is symmetric around its expectation (like the Normal one for instance, the famous symmetric Bell curve), the distribution of \(S\) tends to “lean more to the right” on the horizontal axis i.e. it has a long range of quite large values. This is relevant for us as risk managers, because it is of course exactly those large values that are in particular worrying!
To express this phenomenom mathematically, you want to look at probabilities for \(S\). We have a strategy for computing these in principle (recall from Proposition 7.2) but this would be rather painful indeed to pursue here, so let’s rather grab back to our Monte Carlo simulation technique from Chapter 1 & Chapter 2. In Listing 8.1 below we use the same techniques as in Listing 1.3 and Exercise 1.6. You can see this “long right tail” of \(S\) manifest itself in the histogram shown: the possible values tail off slower and extend further to the right of the red vertical line at level \(\E[S]=7600\) than they do to the left. And the estimated probabilities, showing that \(\P(S>14000)>0.01\) and \(\P(S>15000)<0.01\), mean that if we would use the Value-at-Risk at level \(0.9\) to quantify \(S\)/decide how much to keep in reserve for \(S\) (cf. Section 6.3.1), then this would be an amount somewhere between £14,000 and £15,000!
And, in general this phenomenom could be considerably more pronounced even. In the situation at hand, the common distribution Equation 8.5 of the \(X_i\)’s is helping us out as it has a largest value, keeping somewhat of a bound on how bad things can get. If we would for instance give the \(X_i\)’s a common \(\text{Geometric}(1/760)\) distribution (the parameter chosen so that we still have \(\E[X_1]=760\) and hence also \(\E[S]=7600\)) then the expected loss is unchanged at £7,600 but the lack of an upper bound on the values that the \(X_i\)’s can take causes \(S\) to have even more of a “long right tail”, as Listing 8.2 illustrates (compare this histogram with the previous one, also note the scale difference on the horizontal axis). Note also how the estimated probabilities show that the Value-at-Risk at level \(0.9\) for \(S\) now is even somewhere between £17,000 and £18,000!
Second example — even longer but hopefully in particular even more helpful!!
Example 8.2 Now let’s pick up that setting from Example 8.1 above, with \(S\) given by Equation 8.6, the common distrubution of the \(X_i\)’s by Equation 8.5, and \(N \sim \text{Poisson}(10)\), and let’s mix some risk sharing as described in Proposition 8.1 in there!
Imagine that you enter a risk sharing agreement with company Bark on an excess of loss basis with retention level \(K=600\) (recall from Section 7.2.2) — so you are the cedant and Bark is the excess provider. Bark requires a premium for this service, which they compute by applying the Variance Premium principle (cf. Section 6.2) with loading factor \(\beta=0.005\) applied to the aggregate risk/loss they take on.
So, following what we said in Proposition 8.1, this means that the total risk/loss \(S\) you were originally responsible for, and we analysed in Example 8.1, now gets split up in two parts, say \(S_1\) (your part) and \(S_2\) (Bark’s part) given by \[S_1=\sum_{i=1}^N Y_i \quad \text{and} \quad S_2=\sum_{i=1}^N Z_i,\] where each refund \(X_i\) to be paid to a TV buyer is split up in the parts \[Y_i=\min\{X_i,K\}=\min\{X_i,600\}\] paid by you and \[Z_i=\max\{X_i-K,0\}=\max\{X_i-600,0\}\] paid by Bark (recall that we discussed this type of sharing in detail in Section 7.2.2).
We haven’t seen excess of loss applied to a discrete risk/random variable before, but it is at least as intuitive as for a continuous risk/random variable. Recall from Equation 8.5 that the \(X_i\)’s have range/possible values \(100\), \(500\) and \(1000\), and we can nicely see what the parts paid by you and Bark look like in each of these cases:
- If \(X_i=100\), then \(Y_i=\min\{100,600\}=100\) and \(Z_i=\max\{100-600,0\}=0\) i.e. you pay the whole amount \(100\) and Bark pays nothing;
- If \(X_i=500\), then \(Y_i=\min\{500,600\}=500\) and \(Z_i=\max\{500-600,0\}=0\) i.e. again you pay the whole amount \(500\) and Bark pays nothing;
- If \(X_i=1000\), then \(Y_i=\min\{1000,600\}=600\) and \(Z_i=\max\{1000-600,0\}=400\) i.e. you pay \(400\) while Bark pays \(400\).
The expectation and variance of \(S_1\). Let’s first focus on your situation, now facing the total risk/loss \[S_1=\sum_{i=1}^N Y_i\] and let’s compute the expectation and variance of your total risk/loss. The beauty of it all, as also noted in Proposition 8.2, is that \(S_1\) (still/again) has a compound distribution and so we can again simply apply Theorem 7.1 to be able to write down \[\E[S_1]=\E[N] \E[Y_1] \quad \text{and} \quad \var(S_1)=\var(N) \big( \E[Y_1] \big)^2 + \E[N] \var(Y_1). \tag{8.9}\] As \(N\) is unchanged from Example 8.1, it still has the same \(\text{Poisson}(10)\) distribution with \(\E[N]=10\) and \(\var(N)=10\). So it remains to compute the expectation and variance of \(Y_1\).
There are at least two ways to do this — let’s for clarity stick with the same general strategy we also used before in Example 7.2 for instance: observe that \(Y_1\) is a function of \(X_1\), indeed \(Y_1=f_1(X_1)\) with \(f_1(x)=\min\{x,600\}\), and that we know the distribution of \(X_1\) (cf. Equation 8.5) so that we can apply the standard formula for the expectation of a function of a discrete random variable (cf. Equation A.9): \[\begin{align*} \E[Y_1] &= \E[f_1(X_1)] \\ &= \sum_{k \in R} f_1(k) \P(X_1=k) \\ &= f_1(100) \cdot 0.1 + f_1(500) \cdot 0.3 + f_1(1000) \cdot 0.6 \\ &= 100 \cdot 0.1 +500 \cdot 0.3 +600 \cdot 0.6 \\ &= 520, \end{align*} \] where the second equality uses Equation A.9 with the choice \(h=f_1\) and the third works out the summation using the distribution of \(X_1\) from Equation 8.5.
Pretty much analogue, since squaring both sides of \(Y_1=f_1(X_1)\) gives that \(Y_1^2=f_1(X_1)^2\) we can work out: \[\begin{align*} \E \left[ Y_1^2 \right] &= \E \left[ f_1(X_1)^2 \right] \\ &= \sum_{k \in R} f_1(k)^2 \P(X_1=k) \\ &= f_1(100)^2 \cdot 0.1 + f_1(500)^2 \cdot 0.3 + f_1(1000)^2 \cdot 0.6 \\ &= 100^2 \cdot 0.1 +500^2 \cdot 0.3 +600^2 \cdot 0.6 \\ &= 292000, \end{align*} \] where the only difference is that we now used Equation A.9 with the choice \(h(x)=f_1(x)^2\). Hence \[\var(Y_1)=\E \left[ Y_1^2 \right] - \big( \E[Y_1] \big)^2=292000-520^2=21600.\]
The alternative way for computing the expectation and variance of \(Y_1\) alluded to above would be to use that in this discrete case it is actually quite easy to find out what the distribution of \(Y_1\) itself is — in fact when we inspected the three possible cases earlier in this example we already saw that it looks as follows: \[Y_1 = \begin{cases} 100 & \text{with prob } 0.1 \\ 500 & \text{with prob } 0.3 \\ 600 & \text{with prob } 0.6. \end{cases} \] Computing the expectation and variance directly from this distribution should (of course!) give you the same results as we found above!
To wrap up this computation, if we now plug everything back into Equation 8.9 we arrive at \[\E[S_1]=\E[N] \E[Y_1] =10 \cdot 520=5200\] and \[\begin{align*} \var(S_1) &= \var(N) \big( \E[Y_1] \big)^2 + \E[N] \var(Y_1) \\ &= 10 \cdot 520^2 + 10 \cdot 21600 \\ &= 2920000. \end{align*} \]
It’s nice to observe that when you were still responsible for the whole risk yourself, your aggregate risk/loss had expectation 7600 and variance 6760000, so in this risk sharing situation these have both gone down substantially — as you would expect! For the expectation that is clear, for the variance maybe less so: the argument is that the only effective change is that your largest possible payment per refund has gone down from £1,000 to £600, taking some off the edge of these large payments and causing somewhat less of a spread in the values of the payments per refund and hence also in your aggregate risk/loss.
The expectation and variance of \(S_2\), and Bark’s premium. Let’s now step into the shoes of company Bark! As discussed earlier on, Bark’s aggregate loss/risk is given by \[S_2=\sum_{i=1}^N Z_i,\] where \(Z_i=f_2(X_i)\) with the function \(f_2\) given by \(f_2(x)=\max\{x-600,0\}\).
Again as per Proposition 8.2, \(S_2\) also has a compound distribution and again invoking Theorem 7.1 gives \[\E[S_2]=\E[N] \E[Z_1] \quad \text{and} \quad \var(S_1)=\var(N) \big( \E[Z_1] \big)^2 + \E[N] \var(Z_1).\] Following exactly the same arguments as above, we can now compute \[\begin{align*} \E[Z_1] &= \E[f_2(X_1)] \\ &= \sum_{k \in R} f_2(k) \P(X_1=k) \\ &= f_2(100) \cdot 0.1 + f_2(500) \cdot 0.3 + f_2(1000) \cdot 0.6 \\ &= 0 \cdot 0.1 +0 \cdot 0.3 +400 \cdot 0.6 \\ &= 240 \end{align*} \] and \[\begin{align*} \E \left[ Z_1^2 \right] &= \E \left[ f_2(X_1)^2 \right] \\ &= \sum_{k \in R} f_2(k)^2 \P(X_1=k) \\ &= f_2(100)^2 \cdot 0.1 + f_2(500)^2 \cdot 0.3 + f_2(1000)^2 \cdot 0.6 \\ &= 0^2 \cdot 0.1 +0^2 \cdot 0.3 +400^2 \cdot 0.6 \\ &= 96000, \end{align*} \] so \[\var(Z_1)=\E \left[ Z_1^2 \right] - \big( \E[Z_1] \big)^2=96000-240^2=38400.\]
This brings us to \[\E[S_2]=\E[N] \E[Z_1] =10 \cdot 240=2400\] and \[\begin{align*} \var(S_2) &= \var(N) \big( \E[Z_1] \big)^2 + \E[N] \var(Z_1) \\ &= 10 \cdot 240^2 + 10 \cdot 38400 \\ &= 960000. \end{align*} \]
Now we can also compute the amount of premium that Bark requires to enter this sharing agreement: recall from above that Bark uses the Variance Premium principle with loading factor \(\beta=0.005\), applied to the aggregate risk/loss that they take on i.e. to \(S_2\). That is to say, cf. Section 6.2, they require the amount \[\begin{align*} \pi(S_2) &= \E[S_2] + \beta \var(S_2) \\ &= 2400 + 0.005 \cdot 960000 \\ &= 7200 \end{align*} \] i.e. you’ll need to pay them the handsome sum of £7,200 for them to be willing to enter this risk sharing agreement with you!
A note on linearity. Some of you may have realised that we could have taken a shorter route in some of the comps for Bark above. Indeed fundamentally, every refund payments gets split up between you and Bark i.e. it holds that \[X_i=Y_i+Z_i\] (as is always the case, remember from Section 7.2 — after all one way or the other the refund needs to be paid in total!). This means that by linearity of expectations we have that \[\E[X_1]=\E[Y_1]+\E[Z_1],\] so when we started the Bark computation, since we already had done \(\E[X_1]\) (in Example 8.2) and \(\E[Y_1]\) (earlier in this example) we could immediately have written down what \(\E[Z_1]\) is.
Similarly we could have very quickly found \(\E[S_2]\) by making use of the relationship \(S=S_1+S_2\) (cf. Proposition 8.1) and the fact that we had already computed \(\E[S]\) and \(\E[S_1]\).
However, be aware that we cannot pull this little stunt when it comes to the variances: because \(Y_1\) and \(Z_1\) are (clearly!) not independent and neither are \(S_1\) and \(S_2\), there is no linearity for variances (recall Equation A.55 and that independence a crucial assumption is there!).
The number of non-zero payments that Bark has to do. As a final computation, here is a pretty neat application of Proposition 8.3! :). Recall the three cases that can happen if a refund has to be paid: if the refund amount \(X_i\) is either £100 or £500 then Bark doesn’t have to pay anything. Only if the refund amount \(X_i\) is £1,000 then Bark actually has to pay something (namely £400). It is an interesting question how many payments of £400 Bark then actually has to make! Let’s denote this number by \(\widehat{N}\).
Of course \(\widehat{N}\) will be a random variable, and at first glance the only thing we can really say about it is that it will take a value somewhere in \(\{0,1,\ldots,N\}\), with \(\widehat{N}=0\) meaning that Bark makes no payments at all i.e. that all of the \(N\) refund requests were for amounts less than £1,000, and \(\widehat{N}=N\) meaning that each of the \(N\) refund requests were for the maximal amount of £1,000.
But this is exactly where Proposition 8.3 shines! Indeed it tells us that if we want to count the number of \(X_i\)’s that take the value \(1000\), then we can express this as \[\widehat{N}=\sum_{i=1}^N I_i,\] where \[I_i=\begin{cases} 1 & \text{if } X_i=1000 \\ 0 & \text{if } X_i \not= 1000, \end{cases} \] and in particular this expression shows that \(\widehat{N}\) also has a compound distribution!
And it doesn’t end here — recall that \(N\) has a Poisson distribution? It turns out that \(\widehat{N}\) also has a Poisson distribution! This is very special, there is a priori no reason at all that \(\widehat{N}\) would be in the same distribution family as \(N\)! This phenomenom holds more generally, also in the context of Poisson (point) processes, and is effectively a special case of an operation called Thinning.
Can we prove that this is the case? Yes we can! Recall (from Section A.2) that mgf’s uniquely determine distributions. And mgf’s are convenient to use here, because now we have expressed \(\widehat{N}\) as a compound distribution, we get from Theorem 7.1 immediately the following expression for the mgf \(M_{\widehat{N}}\) of \(\widehat{N}\): \[M_{\widehat{N}}(t)=M_N \big( \log(M_I(t)) \big), \tag{8.10}\] where \(M_N\) is the mgf of \(N\) and \(M_I\) the common mgf of the \(I_i\)’s.
Since \(N \sim \text{Poisson}(10)\), we can readily find in Definition 7.2 (with \(\lambda=10\)) that \[M_N(t)=\exp \left( 10 (e^t-1) \right) \quad \text{for all } t \in \R.\] To get our hands on \(M_I\) we’ll just need to compute it. We follow here exactly the same strategy as we used for the common mgf of the \(X_i\)’s in Example 8.1: fix a \(t \in \R\), write down the def of the mgf from Section A.2, observe that is an expectation and so work this out using the standard formula for the expectation of a function of a discrete random variable (cf. Equation A.9, with the choice \(h(x)=e^{tx}\) and note that the range of \(I_1\) is simply \(R:=\{0,1\}\)): \[\begin{align*} M_I(t) &= \E \left[ e^{t I_1} \right] \\ &= \E \left[ h(I_1) \right] \\ &= \sum_{k \in R} h(k) \P(I_1=k) \\ &= h(0) \cdot \P(I_1=0) + h(1) \cdot \P(I_1=1) \\ &= h(0) \cdot \P(X_1 \not= 1000) + h(1) \cdot \P(X_1=1000) \\ &= 1 \cdot 0.4 + e^t \cdot 0.6 \\ &= 0.6e^t+0.4. \end{align*} \] Now plug things back into Equation 8.10: \[\begin{align*} M_{\widehat{N}}(t) &= M_N \big( \log(M_I(t)) \big) \\ &= \exp \left( 10 \left( e^{\log(M_I(t))}-1 \right) \right) \\ &= \exp \left( 10 \left( M_I(t)-1 \right) \right) \\ &= \exp \left( 10 \left( 0.6e^t+0.4-1 \right) \right) \\ &= \exp \left( 6 \left( e^t-1 \right) \right) \end{align*} \] for all \(t \in \R\). And presto: according to Definition 7.2 this is exactly the mgf of a \(\text{Poisson}(6)\) distribution. And since, as mentioned above, mgf’s uniquely characterise distributions, it follows that in fact \(\widehat{N}\) has a \(\text{Poisson}(6)\) distribution!! Gasps all round, some fainting
Bonus bit ;): these observations give us another way to look at \(S_2\) as well! This case is kind of special: recall that if Bark needs to make a payment, then it is always £400. So if \(\widehat{N}\) gives us the number of payments that Bark needs to do, then the total risk/loss for Bark, i.e. \(S_2\), simply equals \(400 \widehat{N}\). And exploiting our jaw dropping discovery that \(\widehat{N} \sim \text{Poisson}(6)\), it then also immediately follows that \[\E[S_2]=\E[400 \widehat{N}]=400 \E[\widehat{N}]=400 \cdot 6=2400\] and \[\var(S_2)=\var(400 \widehat{N})=400^2 \var(\widehat{N})=400^2 \cdot 6=960000,\] reassuringly the same values as we found earlier! :). Now, this observation really relies on the convenient fact that any payment made by Bark has always the same value, which is not generally the case, so this is not a generally useful insight but nevertheless nice to point out I thought.
Yeah alright already man, blah blah — but is this sharing agreement actually worth it for you?? The million dollar question is of course, after having done all this lengthy analysis, whether it makes good business sense for you to enter this sharing agreement! So you see yourself faces with the following two options:
- Do not enter the agreement. In this case, you remain solely responsible for all refund payments i.e. you’ll have to fully pay whatever value \(S\), as analysed in Example 8.1, will end up taking.
- Do enter the agreement. In this case you will have to pay the required premium of £7,200 to Bark as well as your share of the refund payments i.e. whatever value \(S_1\) ends up taking. So in this scenario your loss is \(L:=7200+S_1\).
So essentially you’ll need to choose between two random losses: either \(S\) or \(L\). Very similar to the situation in Exercise 7.3 for instance. And it is not at all a trivial decision: on the one hand, if only a few claims for refunds end up coming in, then \(S\) will take a pretty small value, easily much smaller than 7200 and hence also than \(L\). On the other hand, if it turns out that many refunds of £1,000 need to be paid then you would regret it if you hadn’t entered the agreement!
As risk manager your main interest will be which option best protects you against large losses. To get some feel for this, see the R code in Listing 8.3 below, which uses the same ideas as in Listing 8.1 & Listing 8.2 to produce a load of samples from \(S\) and \(L\), shows histograms for both as well as does some estimation of Value-at-Risk values (recall from Section 6.3) at levels 0.9, 0.95 and 0.99. Staring at the right tails in the histograms does not seem to reveal much obvious difference between \(S\) and \(L\). Looking at the VaR values, these are clearly in \(S\)’s advantage, so on that basis you would choose not to enter the agreement. Note that in the code you can adjust the value for the loading factor \(\beta\) that Bark uses in its premium calculation, maybe you want to play a bit around with it and see for what loading factor values these VaR values are in \(L\)’s advantage!
8.3 About the sum of aggregate risks/compound distributions
Finally in this block of two chapters on risk sharing and aggregation, we want to turn our attention briefly to the following both mathematically beautiful and practically very useful result. As argued several times before, as professional risk manager you typcially have a diversity of risks from a variety of sources to deal with. The below result shows that if we are dealing with two (or more) different sources, each source with an aggregate risk/total loss modelled by a compound Poisson distribution, then your overall total risk/loss (so the total loss to all sources together) follows again a compound Poisson distribution!
Theorem 8.1 Suppose that \(S^{(1)}\) and \(S^{(2)}\) are two independent compound Poisson distributions as defined in Definition 7.2, with parameter pairs \((\lambda^{(1)},F^{(1)})\) and \((\lambda^{(2)},F^{(2)})\) respectively. That is, \[S^{(1)} =\sum_{i=1}^{N^{(1)}} X^{(1)}_i \quad \text{and} \quad S^{(2)} =\sum_{i=1}^{N^{(2)}} X^{(2)}_i,\] where
- \(N^{(1)} \sim \text{Poisson}(\lambda^{(1)})\),
- \(X^{(1)}_1, X^{(1)}_2, \ldots\) is an i.i.d. sequence of random variables with common cdf \(F^{(1)}\),
- \(N^{(2)} \sim \text{Poisson}(\lambda^{(2)})\),
- \(X^{(2)}_1, X^{(2)}_2, \ldots\) is an i.i.d. sequence of random variables with common cdf \(F^{(2)}\),
- the random variables listed in the above bullet points are all independent of each other.
Define the sum \(S := S^{(1)}+S^{(2)}\). Then \(S\) (also) has a compound Poisson distribution with parameter pair \((\lambda,F)\), where \[\lambda=\lambda^{(1)}+\lambda^{(2)}\] and \[F(x)=\frac{\lambda^{(1)}}{\lambda^{(1)}+\lambda^{(2)}} F^{(1)}(x) + \frac{\lambda^{(2)}}{\lambda^{(1)}+\lambda^{(2)}} F^{(2)}(x) \quad \text{for all } x \in \R. \tag{8.11}\] That is, we can represent \(S\) as \[S=\sum_{i=1}^N X_i, \tag{8.12}\] where
- \(N \sim \text{Poisson}(\lambda^{(1)}+\lambda^{(2)})\);
- \(X_1, X_2, \ldots\) is an i.i.d. sequence of random variables, independent of \(N\), whose common cdf \(F\) is given by Equation 8.11.
Some notes:
If you feel slightly uncomfortable with that “we can represent as” statement, then note that you can read it as follows: for the purpose of computing any probability/expectation for \(S\), you can simply see \(S\) as equal to/given by Equation 8.12 (so that you can for instance apply all techniques/results from Section 7.3.2 to \(S\)!). Technically the statement is that \(S\) has the same distribution as the expression in the right hand side of Equation 8.12.
Note that we really need compound Poisson distributions here, it doesn’t say anything about other compound distributions — indeed this result is generally not true for other compound distributions!
This is a situation in which the principle “if it applies to two, it also applies to \(n \geq 3\) of them” holds! Indeed if we for instance were to have three independent compound Poisson distributed \(S^{(1)}, S^{(2)}, S^{(3)}\), then their sum \(S=S^{(1)}+S^{(2)}+S^{(3)}\) is still a compound Poisson distribution: just apply this result first to \(\widetilde{S}:=S^{(1)}+S^{(2)}\), and then again to \(S=\widetilde{S}+S^{(3)}\).
This is a very elegant and pretty result I think! :). Intuitively it tells us the following. If you are managing two sources of loss, say source 1 and source 2, with compound Poisson distributed aggregate risks/losses \(S^{(1)}\) and \(S^{(2)}\) resp., then you can equally well see this as one single source of loss, whose aggregate loss \(S\) has the following characteristics. The number of losses appearing (\(N\)) is simply the sum of the number of losses from source 1 (\(N^{(1)}\)) and the number of losses from source 2 (\(N^{(2)}\)). The size of each loss, \(X_i\), is constructed (in Equation 8.11) as the properly weighted average of the loss sizes from each individual source, where the weighting accounts for the fact that one source may contribute more losses than the other. Essentially \(X_i\) is the answer to the question: “if I have \(N^{(1)}\) losses from source 1 and \(N^{(2)}\) losses from source 2, and we would pick a loss at random from these \(N^{(1)}+N^{(2)}\) ones, what size would we expect it to have?”.
Compare it with the following: if we have a sequence of \(5\) numbers all equal to \(10\) (source 1) and a sequence of \(20\) numbers all equal to \(2\) (source 2), then in total we have the same as when we had one single sequence of \(5+20=25\) numbers, all equal to a properly weighted average of the numbers from these sources, namely (the analogue of Equation 8.11) \[\frac{5}{25} \cdot 10 + \left( 1-\frac{5}{25} \right) \cdot 2=3.6.\] Indeed, this one single source yields a total of \(25 \cdot 3.6=90\), while adding up the numbers from sources 1 and 2 seperately gives \(5 \cdot 10+20 \cdot 2=90\) as well!
Proof. How do we prove something like this? Well we need to show that the two random variables \[S:=S^{(1)}+S^{(2)}=\sum_{i=1}^{N^{(1)}} X^{(1)}_i + \sum_{i=1}^{N^{(2)}} X^{(2)}_i \tag{8.13}\] and \[\widetilde{S}:=\sum_{i=1}^N X_i \tag{8.14}\] have the same distribution. The standard way of doing that is showing that they have the same cdf’s. But that’s rather inconvenient in this case — for the cdf we need to compute probabilities, and for these we only have the strategy from Proposition 7.2 available, plus that Equation 8.13 is the sum of two compound distributions rather than a single one. Bluh.
But what about mgf’s? We know that mgf’s uniquely identify distributions (cf. Section A.2), so it would be sufficient to show that \(S\) and \(\widetilde{S}\) have the same mgf. Further we have a nice compact expression for the mgf of compound distributions from Theorem 7.1, and the fact that in Equation 8.13 we have the sum of two compound distributions is not too much of a problem as these are independent and independent sums work very nicely with mgf’s (cf. Equation A.19). So let’s go down this route!
Firstly, combining Equation A.19 with Theorem 7.1 we can work out the mgf of \(S\) as follows: \[ \begin{align*} M_S(t) &= M_{S^{(1)}}(t) M_{S^{(2)}}(t) \\ &= M_{N^{(1)}} \Big( \log \big( M_{X^{(1)}}(t) \big) \Big) M_{N^{(2)}} \Big( \log \big( M_{X^{(1)}}(t) \big) \Big) \\ &= \exp \bigg( \lambda^{(1)} \Big( \exp \Big( \log \big( M_{X^{(1)}}(t) \big) \Big) -1 \Big) \bigg) \\ & \qquad \cdot \exp \bigg( \lambda^{(2)} \Big( \exp \Big( \log \big( M_{X^{(2)}}(t) \big) \Big) -1 \Big) \bigg) \\ &= \exp \big( \lambda^{(1)} (M_{X^{(1)}}(t)-1)+\lambda^{(2)} (M_{X^{(2)}}(t)-1) \big), \end{align*} \] where we also used the expression for the mgf of a Poisson distribution we have seen in Definition 7.2.
Secondly, applying Theorem 7.1 to Equation 8.14 and again using the well known expression for the mfg of a Poisson distribution: \[ \begin{align*} M_{\widetilde{S}}(t) &= M_{N} \Big( \log \big( M_X(t) \big) \Big) \\ &= \exp \bigg( (\lambda^{(1)}+\lambda^{(2)}) \Big( \exp \Big( \log \big( M_X(t) \big) \Big) -1 \Big) \bigg) \\ &= \exp \big( (\lambda^{(1)}+\lambda^{(2)}) (M_X(t)-1) \big). \end{align*} \] So in order to arrive at our desired conclusion that \(M_S=M_{\widetilde{S}}\) it remains to show that \[(\lambda^{(1)}+\lambda^{(2)}) (M_X(t)-1)=\lambda^{(1)} \left( M_{X^{(1)}}(t)-1 \right)+\lambda_2 \left( M_{X^{(2)}}(t)-1 \right)\] i.e. that \[M_X(t) = \frac{\lambda^{(1)}}{\lambda^{(1)}+\lambda^{(2)}} M_{X^{(1)}}(t) + \frac{\lambda^{(2)}}{\lambda^{(1)}+\lambda^{(2)}} M_{X^{(2)}}(t). \tag{8.15}\] So what do we know about the \(X_i\)’s that we could use here? Well all we have is the expression for its cdf from Equation 8.11 so we should be looking to put that one in play here. With slightly more advanced Prob tools than we have available here this is a pretty easy exercise. An outline for a way to do this with our available tools (don’t worry about it for the exam of course):
- Define two new random variables, independent of all of the ones we have seen so far and of each other: \(U \sim \text{Unif}(0,1)\) and \(\widetilde{X}\) given by \[\widetilde{X}=\begin{cases} X^{(1)}_1 & \text{if } U \leq \lambda^{(1)}/(\lambda^{(1)}+\lambda^{(2)}) \\ X^{(2)}_1 & \text{if } U > \lambda^{(1)}/(\lambda^{(1)}+\lambda^{(2)}) \end{cases} \tag{8.16}\]
- Use the Law of Total Probability (cf. Equation A.46) to show that the cdf \(F_{\widetilde{X}}\) of \(\widetilde{X}\) is given by \[F_{\widetilde{X}}(x)=\frac{\lambda^{(1)}}{\lambda^{(1)}+\lambda^{(2)}} F^{(1)}(x) + \frac{\lambda^{(2)}}{\lambda^{(1)}+\lambda^{(2)}} F^{(2)}(x) \quad \text{for all } x \in \R.\] As this is identical to the cdf of \(X\) (cf. Equation 8.11) and cdf’s uniquely determine distributions, \(X\) and \(\widetilde{X}\) have the same distribution. And since mgf’s also uniquely determine distributions (cf. Section A.2), they also have the same mgf’s. So, to prove Equation 8.15 it is sufficient to show that \[M_{\widetilde{X}}(t) = \frac{\lambda^{(1)}}{\lambda^{(1)}+\lambda^{(2)}} M_{X^{(1)}}(t) + \frac{\lambda^{(2)}}{\lambda^{(1)}+\lambda^{(2)}} M_{X^{(2)}}(t). \tag{8.17}\]
- Use the Law of Total Expectation (cf. Equation A.47) to work out \(M_{\widetilde{X}}\) (from Equation 8.16) and show that Equation 8.17 indeed holds.
Alternatively, if you’re willing to make a sacrifice when it comes to our assumptions, you could add the assumption that the \(X^{(1)}_i\)’s and \(X^{(2)}_i\)’s are continuous random variables. Then Equation 8.11 shows that the \(X_i\)’s are also continuous random variables, and differentiating Equation 8.11 gives the following relationship between the common pdf’s of these three sequences of random variables: \[f_X(x)=\frac{\lambda^{(1)}}{\lambda^{(1)}+\lambda^{(2)}} f_{X^{(1)}}(x) + \frac{\lambda^{(2)}}{\lambda^{(1)}+\lambda^{(2)}} f_{X^{(2)}}(x) \quad \text{for all } x \in \R.\] If you now express the mgf’s appearing in Equation 8.15 as integrals (using Equation A.16), then this relationship between the pdf’s immediately yields Equation 8.15.
Let’s see some of this in action!
Example 8.3 Suppose that you are responsible for two sources of risk, say source A and source B, with the following characteristics:
Source A: your aggragate/total risk/loss \(S^{(A)}\) can be expressed as \[S^{(A)}=\sum_{i=1}^{N^{(A)}} X_i^{(A)},\] where \(N^{(A)} \sim \text{Poisson}(5)\) and \(X_1^{(A)}, X_2^{(A)}, \ldots\) is a sequence of i.i.d. random variables, independent of \(N^{(A)}\), with a common \(\text{Unif}(0,10)\) distribution.
So \(S^{(A)}\) has a compound Poisson distribution (cf. Definition 7.2) with parameter pair \((5,F^{(A)})\), where the common cdf \(F^{(A)}\) of \(X_1^{(A)}, X_2^{(A)}, \ldots\) is given by \[F^{(A)}(x)=\begin{cases} 0 & \text{if } x<0 \\ x/10 & \text{if } x \in [0,10) \\ 1 & \text{if } x \geq 10 \end{cases} \] (if you don’t remember this one, never forget that you can find the relevant pdf in Appendix B and then obtain the corresponding cdf via Equation A.12).
Source B: your aggragate/total risk/loss \(S^{(B)}\) can be expressed as \[S^{(B)}=\sum_{i=1}^{N^{(B)}} X_i^{(B)},\] where \(N^{(B)} \sim \text{Poisson}(10)\) and \(X_1^{(B)}, X_2^{(B)}, \ldots\) is a sequence of i.i.d. random variables, independent of \(N^{(B)}\), with a common \(\text{Unif}(0,5)\) distribution.
So \(S^{(B)}\) also has a compound Poisson distribution with parameter pair \((10,F^{(B)})\), where the common cdf \(F^{(B)}\) of \(X_1^{(B)}, X_2^{(B)}, \ldots\) is given by \[F^{(B)}(x)=\begin{cases} 0 & \text{if } x<0 \\ x/5 & \text{if } x \in [0,5) \\ 1 & \text{if } x \geq 5 \end{cases} \] (if you don’t remember this one, never forget that you can find the relevant pdf in Appendix B and then obtain the corresponding cdf via Equation A.12).
Then your aggregate/total risk \(S\) for both sources together, \[S:=S^{(A)}+S^{(B)} \tag{8.18}\] has according to Theorem 8.1 again a compound Poisson distribution with parameter pair \((15,F)\) where the cdf \(F\) is given by \[F(x)=\frac{5}{15} F^{(A)}(x) + \frac{10}{15} F^{(B)}(x) \quad \text{for all } x \in \R.\] Plugging in the above expressions for \(F^{(A)}\) and \(F^{(B)}\) allows us to easily work out \[F(x)=\begin{cases} 0 & \text{if } x<0 \\ x/6 & \text{if } x \in [0,5) \\ x/30+2/3 & \text{if } x \in [5,10) \\ 1 & \text{if } x \geq 10. \end{cases} \tag{8.19}\]
So the meaning of the statement in Theorem 8.1 is that the random variable \(S\), defined as in Equation 8.18, has the same distribution as the random variable \[\widetilde{S}:=\sum_{i=1}^N X_i, \tag{8.20}\] where \(N \sim \text{Poisson}(15)\) and \(X_1, X_2, \ldots\) is a sequence of i.i.d. random variables, independent of \(N\), with common cdf \(F\) given by Equation 8.19. Practically that means that any probability, expectation, mgf, … you want to compute for \(S\) may equivalently be computed for \(\widetilde{S}\), and since this guy has a compound (Poisson) distribution we can throw all our tools from Section 7.3.2 at it! :).
Finally then, as a quick and dirty visual confirmation, let’s produce a bunch of samples from \(S\) as defined in Equation 8.18, a bunch of samples from \(\widetilde{S}\) as defined in Equation 8.20, and then let’s compare the histograms of these samples. The maths tells us that their distributions are exactly the same, and so we should expect that these histograms look at least very similar (allowing for some error because, as we know from our earlier discussions around Monte Carlo simulations in Chapter 1 & Chapter 2, such a sample based approach gives only an approximation of the underlying distribution). See Listing 8.4 below. We use the same general ideas as in the code blocks Listing 8.1–Listing 8.3 we did earlier on in this chapter, with one additional note. To produce samples from \(\widetilde{S}\), we need to produce samples from the \(X_i\)’s. From these guys all we know is their common cdf, cf. Equation 8.19. This is not a “well known” distribution built into R, and hence we need to reach back to one of our algorithms discussed in Chapter 2. Since we have the cdf and it’s a pretty simple and easily invertible expression, our old friend the inverse transform method from Theorem 2.1 comes in very handy! :).
8.4 Further reading
This use of compound distributions is part of a much wider mathematical field around the modelling of aggregate risks/losses, involving Poisson (point) processes/renewal processes, Ruin Theory/Cramér-Lundberg processes and more general Lévy processes etc. It is commonly associated with insurance, where it definitely plays an important role, but as we have highlighted there are many other important applications as well. If you’re interested in more, see e.g. Chapters 2, 3 and 4 in Kaas et al. (2008) or Mikosch (2004).
8.5 Some exercises
Each exercise has a (rough) indication of its difficulty, as follows:
| * | easier: can be solved by (almost) only using relevant definitions/results, |
| ** | medium: in addition to relevant definitions/results, needs a limited amount of work/creativity, |
| *** | harder: in addition to relevant definitions/results, needs a larger amount of work/serious creativity, |
| 💀 | warning: might make your brain hurt! These are mainly intended to provide some extra challenge for those of you keen on that and are generally quite hard. You don't need to worry about these too much for exam purposes. |
The exam consists of mostly ** and *** level questions, some *, and possibly at most a few marks worth of 💀.
A bit of preaching: it is an incredibly important part of the study process to try and work on the exercises as much as possible. To become a better mathematician/learn new maths (and also to get a good exam mark ;)), above all you need to do it. And yes, of course that includes falling over things, and making mistakes, and getting stuck, and getting frustrated — all part of the game and what you’re supposed to be doing! Your lecturers have done that as well and still do it. What matters is that you don’t let that discourage you and that you make good use of the help and resources available to help you develop your skills. As part of that, many exercises have a hint in a block like this:
Hint!
These are trying to help you on your way if you don’t know where to start or to provide some ideas if you get stuck. In spirit of the above, always have a look at these first and try again before you look at the full solution. (These hints are an extra service that won’t be available in the exam I’m afraid ;).)
Of course, we have our classes and there’s office hours, email etc. as well — I’m at any time very happy to help you with any questions you may have, and you should please never feel that any question is “too dumb” to ask!
Full/detailed solutions for the exercises will become available, just immediately below the exercises, immediately after our tutorial hour (you may have to refresh the page).
Exercise 8.1 [**] Suppose that you are responsible for two independent and identical potential risks. Each of these risks materialises with probability \(1/4\). If a risk materialises, then the loss it results in for you is modelled by a \(\text{Gamma}(2,\alpha)\) distribution, for some \(\alpha>0\).
You denote your aggregate risk/loss by \(S\) and express it as a compound distribution as defined in Definition 7.1 i.e. \[S=\sum_{i=1}^N X_i,\] where \(N\) denotes the number of risks that materialises and \(X_i\) the \(i\)-th loss that you suffer due to a risk materialising.
What is the probability that you suffer no loss at all?
What distribution does \(N\) have? And the \(X_i\)’s?
Compute the expectation and variance of \(S\).
Show that \(X_1+X_2\) has a pdf \(f_{X_1+X_2}\) given by \[f_{X_1+X_2}(x)=\begin{cases} \frac{\alpha^4}{6} x^3 e^{-\alpha x} & \text{if } x>0 \\ 0 & \text{if } x \leq 0. \end{cases} \]
Hint: there is a quick and a not-so-quick way! ;).
Show that for any \(x>0\), the probability that your total loss exceeds \(x\) is given by \[e^{-\alpha x} \left( \frac{1}{96} \alpha^3 x^3 + \frac{1}{32} \alpha^2 x^2 + \frac{7}{16} \alpha x + \frac{7}{16} \right).\]
Hint: computationally a bit annoying, you may first want to use integration by parts (recall that we also used it in Example 6.1 e.g.) to establish that for any \(k \in \{1,2,\ldots\}\) we have that \[\int_x^\infty y^k e^{-\alpha y} \d y = \frac{1}{\alpha} x^k e^{-\alpha x} + \frac{k}{\alpha} \int_x^\infty y^{k-1} e^{-\alpha y} \d y. \tag{8.21}\]
Note that this question is in particular about the material discussed in Section 7.3 and Example 8.1.
For part i, convince yourself that in order to suffer no loss at all, neither of the two risks should materialise.
For part ii, the common distribution of the \(X_i\)’s follows immediately from the question. For the distribution of \(N\) you can either argue directly (along the same lines as in part i) or you can recognise that we have discussed this structure in general before, see Section 7.3.1 e.g.
For part iii, never forget about our treasure trove of results for compound distributions from Section 7.3.2 — in this case in particular Theorem 7.1! Together with Appendix B it’s a straightforward comp!
For part iv, with our knowledge from previous Prob courses this is a one-liner — if you have forgotten, take a quick look at Section A.6!
For part v, observe that we’re asked to compute \(\P(S>x)\). As in part iii, we don’t have much choice but to dig through Section 7.3.2 to find something we could use here. Up pops Proposition 7.2 i, laying out our standard plan of attack for computing a probability for a compound distribution. All you otherwise need is to realise that the probs \(\P(X_1>x)\) and \(\P(X_1+X_2>x)\) can be worked out as integrals involving their pdf’s (recall from Equation A.15), and some elbow grease to make your way through these integrals — use Equation 8.21 as much as you can!
Note that this question is in particular about the material discussed in Section 7.3 and Example 8.1.
For part i, as also discussed in Example 8.1, note that a \(\text{Gamma}(2,\alpha)\) distribution has range \((0,\infty)\) i.e. will always take a strictly positive value (cf. Equation A.18 and if needs be review the pdf of this distribution on Appendix B). So you suffer no loss at all if and only if \(N=0\) i.e. if and only if neither of the two risks materialises. By independence this happens with probability \[\frac{3}{4} \cdot \frac{3}{4}=\frac{9}{16}.\]
For part ii, first for the distribution of \(N\). You could go one of two ways here. The first one is a direct argument. Since there are two risks in total, each of which may or may not materialise, the number of risks that materialises \(N\) has as range/set of possible values \(\{0,1,2\}\). In part i we argued already that \(\P(N=0)=9/16\), and similarly \[\P(N=2)=\frac{1}{4} \cdot \frac{1}{4}=\frac{1}{16}\] and then \[\P(N=1)=1-\P(N=0)-\P(N=1)=1-\frac{9}{16}-\frac{1}{16}=\frac{6}{16}.\]
The second way is to realise that this structure of having \(2\) risks in total that each independently materialises with prob \(1/4\) is one that we discussed in general before, e.g. in Section 7.3.1, and as concluded there we then have \(N \sim \text{Binomial}(2,1/4)\). Obviously this gives exactly the same range and probs as we computed via the direct way, convince yourself using Appendix B if needs be!
Further, it is clear from the question that the \(X_i\)’s have a common \(\text{Gamma}(2,\alpha)\) distribution.
For part iii, as we’ve seen several times before, also in the examples in this chapter: whenever you need to compute any prob/expectation for a compound distribution, you head over to Section 7.3.2! In particular, we know from Theorem 7.1 that \[\E[S]=\E[N] \E[X_1]\] and \[\var(S)=\var(N) \big( \E[X_1] \big)^2 + \E[N] \var(X_1).\]
Working this out is pretty straightforward: we saw in part ii that \(N \sim \text{Binomial}(2,1/4)\) and hence, using Appendix B if needed, \[\E[N]=2 \cdot \frac{1}{4}=\frac{1}{2} \quad \text{and} \quad \var(N)=2 \cdot \frac{1}{4} \cdot \frac{3}{4}=\frac{6}{16}\] (note that if in part ii you used the direct way and didn’t recognise the Binomial distribution, you could of course still work out the expectation and variance of \(N\) using the standard approach to do that for discrete random variables i.e. using Equation A.10 & Equation A.9). Further, since \(X_1 \sim \text{Gamma}(2,\alpha)\) as discussed in part ii, we can immediately read from Appendix B that \[\E[X_1]=\frac{2}{\alpha} \quad \text{and} \quad \var(X_1)=\frac{2}{\alpha^2}.\]
So altogether we arrive at \[\E[S]=\E[N] \E[X_1]=\frac{1}{2} \cdot \frac{2}{\alpha}=\frac{1}{\alpha}\] and \[\begin{align*} \var(S) &= \var(N) \big( \E[X_1] \big)^2 + \E[N] \var(X_1) \\ &= \frac{6}{16} \cdot \left( \frac{2}{\alpha} \right)^2 + \frac{1}{2} \cdot \frac{2}{\alpha^2} \\ &= \frac{5}{2\alpha^2}. \end{align*} \]
For part iv, the quick way is to realise that since \(X_1\) and \(X_2\) both have a \(\text{Gamma}(2,\alpha)\) distribution and they are independent, we know very well from our earlier Prob courses what distribution \(X_1+X_2\) has — indeed a \(\text{Gamma}(4,\alpha)\) distribution (cf. point 2 in Section A.6)! So it is just a matter of writing down the pdf of this guy, for which you could use Appendix B if needs be.
The not-so-quick way would be to go down the convolution route, see e.g. point 3 in Section A.6…
For part v, let’s first of all establish Equation 8.21. As mentioned, using integration by parts as also used in Example 6.1 e.g. we get \[\begin{align*} \int_x^\infty y^k e^{-\alpha y} \d y &= \left. y^k \frac{-1}{\alpha} e^{-\alpha y} \right|_{x}^\infty - \int_x^\infty k y^{k-1} \frac{-1}{\alpha} e^{-\alpha y} \d y \\ &= \frac{1}{\alpha} x^k e^{-\alpha x} + \frac{k}{\alpha} \int_x^\infty y^{k-1} e^{-\alpha y} \d y. \end{align*} \]
Now let’s attack the problem. Note that we are asked to compute \(\P(S>x)\), for some (fixed) \(x>0\). As in part iii, we need to dig through Section 7.3.2 to find a suitable tool. In this case Proposition 7.2 is our friend, in particular part i of it as that is about computing probs for compound distributions. Recalling from part ii that \(\text{Range}(N)=\{0,1,2\}\) and choosing \(B=(x,\infty)\) — after all, \(\P(S>x)=\P(S \in (x,\infty))\) — we get from Proposition 7.2 i that \[\P(S>x)=\sum_{n \in \{0,1,2\}} \P(S>x \, | \, N=n) \P(N=n), \tag{8.22}\] where:
- For \(n=0\), looking at \(\P(S>x \, | \, N=0)\), as Proposition 7.2 tells us this is either \(0\) or \(1\) depending on whether \(0 \in B\). Since we have here \(B=(x,\infty)\) and \(x>0\), we see that \(0 \not\in B\) and hence \[\P(S>x \, | \, N=0)=0.\]
- Next for \(n=1\) i.e. looking at \(\P(S>x \, | \, N=1)\). Proposition 7.2 tells us that this simplifies to \(\P(X_1>x)\), and since \(X_1 \sim \text{Gamma}(2,\alpha)\) we can work this out using an integral involving the pdf of \(X_1\) (which you can again pick up from Appendix B), cf. Equation A.15: \[\begin{align*} \P(X_1>x) &= \int_x^\infty f_{X_1}(y) \d y \\ &= \int_x^\infty \alpha^2 y e^{-\alpha y} \d y \\ &= \alpha x e^{-\alpha x} + \alpha \int_x^\infty e^{-\alpha y} \d y \\ &= \alpha x e^{-\alpha x}+e^{-\alpha x}, \end{align*} \] where the third equality uses Equation 8.21 with \(k=1\). So \[\P(S>x \, | \, N=1)=\alpha x e^{-\alpha x}+e^{-\alpha x}.\]
- Finally for \(n=2\) i.e. looking at \(\P(S>x \, | \, N=2)\). Proposition 7.2 now tells us that this equals \(\P(X_1+X_2>x)\), and since we discussed the pdf of \(X_1+X_2\) in part iv above already we can in similar vein work this out using an integral: \[\begin{align*} \P(X_1+X_2>x) &= \int_x^\infty f_{X_1+X_2}(y) \d y \\ &= \int_x^\infty \frac{\alpha^4}{6} y^3 e^{-\alpha y} \d y \\ &= \frac{\alpha^3}{6} x^3 e^{-\alpha x} + \frac{\alpha^3}{2} \int_x^\infty y^2 e^{-\alpha y} \d y \\ &= \frac{\alpha^3}{6} x^3 e^{-\alpha x} + \frac{\alpha^2}{2} x^2 e^{-\alpha x} + \alpha^2 \int_x^\infty y e^{-\alpha y} \d y \\ &= \frac{\alpha^3}{6} x^3 e^{-\alpha x} + \frac{\alpha^2}{2} x^2 e^{-\alpha x} + \alpha x e^{-\alpha x} + e^{-\alpha x}, \end{align*} \] where the third, fourth and fifth equalities uses Equation 8.21 with \(k=3\), \(k=2\) and \(k=1\) resp. I know! However we got there: \[\P(S>x \, | \, N=2)=\frac{\alpha^3}{6} x^3 e^{-\alpha x} + \frac{\alpha^2}{2} x^2 e^{-\alpha x} + \alpha x e^{-\alpha x} + e^{-\alpha x}.\]
We can now put all the pieces together, also using the probs for \(N\) we already established in part ii, we ultimately get from Equation 8.22 that \[\begin{align*} \P(S>x) &= \sum_{n \in \{0,1,2\}} \P(S>x \, | \, N=n) \P(N=n) \\ &= 0 \cdot \frac{9}{16} + \left( \alpha x e^{-\alpha x}+e^{-\alpha x} \right) \cdot \frac{6}{16} \\ & \qquad + \left( \frac{\alpha^3}{6} x^3 e^{-\alpha x} + \frac{\alpha^2}{2} x^2 e^{-\alpha x} + \alpha x e^{-\alpha x} + e^{-\alpha x} \right) \cdot \frac{1}{16} \\ &= e^{-\alpha x} \left( \frac{1}{96} \alpha^3 x^3 + \frac{1}{32} \alpha^2 x^2 + \frac{7}{16} \alpha x + \frac{7}{16} \right). \end{align*} \]
Exercise 8.2 [**/***] Consider again the collection of two risks from Exercise 8.1 you are responsible for. Suppose now that an external party is offering their risk sharing services to you. They are offering three possible sharing agreements, a, b and c, as listed below. Accepting any of the agreements requires a premium payment from you, which is computed using the Variance Premium principle with loading factor \(\beta>0\) (cf. Section 6.2.1) applied to the total/aggregate risk the external party becomes responsible for. Your job is to compute that premium for each of the three agreements.
- The external party fully pays each risk that materialises.
- A proportional risk sharing agreement with retention level \(\gamma \in (0,1)\) is applied to each risk that materialises. So the external party pays \(100(1-\gamma)\%\) of each risk that materialises.
- An excess loss risk sharing agreement with retention level \(K>0\) is applied to each risk that materialises. So the external party is the excess provider for each risk that materialises.
Hint: the third agreement requires a good bit of elbow grease, sorry — don’t forget about Equation 8.21!
For a general discussion of sharing in the context of a collection of risks, recall Section 8.2 and in particular Proposition 8.1.
For agreements a and b, if you think about it is in both cases actually pretty straightforward to find a simple relationship between \(S_2\), the total risk/loss for the external party, and \(S\) (the total risk/loss for the whole collection that you are responsible for if you don’t share and we analysed in Exercise 8.1). Using that relationship, and the comps already done in Exercise 8.1, the premium is pretty easy to work out!
Agreement c is much more annoying though. First remind yourself of the exact setup here (as also seen in Example 8.2 for instance): we can express \(S_2\) as a compound distribution of the form \[S_2 = \sum_{i=1}^N Z_i\] where \(Z_i=\max\{X_i-K,0\}\). Hence we can appeal to Theorem 7.1 for expressions for the expectation and variance of \(S_2\), and in order to work out the ingredients we need there use the same strategy as also used in Example 8.2 as well as in several exercises in Section 7.4 e.g.: since \(Z_1\) is a function of \(X_1\) and we know the pdf of \(X_1\), we can apply Equation A.16. Quite a bit of algebra to be worked out carefully here!
Note that we are in the context of Section 8.2 here, with the process as laid out in Proposition 8.1: each risk that materialises with loss \(X_i\) gets split up in a part \(Y_i\) to be paid by you and a part \(Z_i\) to be paid by the other party — how the split exactly happens depends on the details in each of these three possible agreements. Then the total/aggregate risk the external party becomes responsible for is given by the compound distribution \[S_2 = \sum_{i=1}^N Z_i. \tag{8.23}\] The premium they require for this agreement, computed using the Variance Premium principle with loading factor \(\beta>0\) (cf. Section 6.2.1) applied to \(S_2\), is hence given by \[\pi(S_2)=\E[S_2]+\beta \var(S_2).\] So let’s work out this \(\pi(S_2)\) for each possible agreement!
Recall that in Exercise 8.1 we already did some work with \(S\), the aggregate risk/loss for this collection.
Agreement a. In this case the external party takes over the whole collection, so logically their aggregate risk loss is just \(S\) i.e. \(S_2=S\). Or, slightly more mathematically: each loss \(X_i\) is split into the trivial parts \(Z_i=X_i\) and \(Y_i=0\), and hence, from Equation 8.23 and Exercise 8.1: \[S_2 = \sum_{i=1}^N Z_i=\sum_{i=1}^N X_i=S.\] Using the work we already did in part iii of Exercise 8.1 we then immediately get that \[\pi(S_2)=\E[S_2]+\beta \var(S_2)=\E[S]+\beta \var(S)=\frac{1}{\alpha}+\frac{5\beta}{2\alpha^2}.\]
Agreement b. Recall the proportional risk sharing agreement also from Section 7.2.1 and Exercise 7.1 for instance: each loss \(X_i\) is split into the parts \(Z_i=(1-\gamma) X_i\) and \(Y_i=\gamma X_i\). Plugging this into Equation 8.23 we see that we can again easily relate \(S_2\) to \(S\): \[S_2 = \sum_{i=1}^N Z_i = \sum_{i=1}^N (1-\gamma) X_i=(1-\gamma) \sum_{i=1}^N X_i=(1-\gamma) S.\] Again, it’s also just common sense: if the external party pays a percentage of each loss, then in total they end up paying that same percentage of the total loss no.
This means that we can again pretty quickly compute \(\pi(S_2)\) using the work we did in part iii of Exercise 8.1 already (don’t slip over the square if you take a multiplicative constant out of a variance eh!): \[\begin{align*} \pi(S_2) &= \E[S_2]+\beta \var(S_2) \\ &=\E[(1-\gamma) S]+\beta \var((1-\gamma) S) \\ &= (1-\gamma) \E[S]+ \beta (1-\gamma)^2 \var(S) \\ &= (1-\gamma) \cdot \frac{1}{\alpha} + \beta (1-\gamma)^2 \cdot \frac{5}{2\alpha^2} \\ &= \frac{1-\gamma}{\alpha} + \frac{5 \beta (1-\gamma)^2}{2\alpha^2}. \end{align*} \]
Agreement c. Also the excess of loss risk sharing agreement we have seen before, several times even: for instance in Section 7.2.2, several exercises in Section 7.4, and in this chapter in Example 8.2. In this case we don’t have such a nice and easily epxloitable relationship between \(S_2\) and \(S\) anymore, so we have to work a bit harder!
Recall (from any of the above mentioned earlier appearances) that in the case of excess of loss sharing with retention level \(K>0\), each loss \(X_i\) is split up into the parts \(Y_i=\min\{X_i,K\}\) and \(Z_i=\max\{X_i-K,0\}\). What we need to be able to compute \(\pi(S_2)\) is the expectation and variance of \(S_2\). For this we can follow the same strategy as we also used in Example 8.2: observe that per Proposition 8.2, \(S_2\) as given by Equation 8.23 has a compound distribution and hence we have again all tools from Section 7.3.2 at our disposal, in particular our maybe most favourite friend Theorem 7.1 which gives us that \[\E[S_2]=\E[N] \E[Z_1] \quad \text{and} \quad \var(S_1)=\var(N) \big( \E[Z_1] \big)^2 + \E[N] \var(Z_1). \tag{8.24}\] In part iii of Exercise 8.1 we already saw that \[\E[N]=\frac{1}{2} \quad \text{and} \quad \var(N)=\frac{6}{16},\] so on to the expectation and variance of \(Z_1\). We use the same idea as we used e.g. in Example 8.2 as well (except that we’re now dealing with a continuous random variable \(X_1\) rather than a discrete one as we did there): since \(Z_1\) is a function of \(X_1\), in particular \(Z_1=f_2(X_1)\) with \(f_2(x)=\max\{x-K,0\}\), and we know the distribution of \(X_1\) (namely \(\text{Gamma}(2,\alpha)\), cf. Exercise 8.1 part ii), we can use the standard formula for the expectation of a function of a continuous random variable i.e. Equation A.16, first with the choice \(h(x)=f_2(x)=\max\{x-K,0\}\) and then with the choice \(h(x)=f_2(x)^2=\max\{x-K,0\}^2\), to compute (where \(f_{X_1}\) denotes the pdf of \(X_1\)) \[\begin{align*} \E[Z_1] &= \E[f_2(X_1)] \\ &= \int_{-\infty}^\infty f_2(x) f_{X_1}(x) \d x \\ &= \int_0^\infty \max\{x-K,0\} \alpha^2 x e^{-\alpha x} \d x \\ &= \int_K^\infty (x-K) \alpha^2 x e^{-\alpha x} \d x \\ &= \int_K^\infty \alpha^2 x^2 e^{-\alpha x} \d x-K \int_K^\infty \alpha^2 x e^{-\alpha x} \d x \\ &= \left( K+\frac{2}{\alpha} \right) e^{-\alpha K}, \end{align*} \] where in the final equality we made repeated use of Equation 8.21, and \[\begin{align*} \E[Z_1^2] &= \E[f_2(X_1)^2] \\ &= \int_{-\infty}^\infty f_2(x)^2 f_{X_1}(x) \d x \\ &= \int_0^\infty \max\{x-K,0\}^2 \alpha^2 x e^{-\alpha x} \d x \\ &= \int_K^\infty (x-K)^2 \alpha^2 x e^{-\alpha x} \d x \\ &= \int_K^\infty \alpha^2 x^3 e^{-\alpha x} \d x-2K \int_K^\infty \alpha^2 x^2 e^{-\alpha x} \d x + K^2 \int_K^\infty \alpha^2 x e^{-\alpha x} \d x \\ &= \left( \frac{2K}{\alpha}+\frac{6}{\alpha^2} \right) e^{-\alpha K}, \end{align*} \] again making good use of Equation 8.21 for the final step.
So we get that \[\var(Z_1)=\E[Z_1^2]-\big( \E[Z_1] \big)^2=\left( \frac{2K}{\alpha}+\frac{6}{\alpha^2} \right) e^{-\alpha K} - \left( K+\frac{2}{\alpha} \right)^2 e^{-2 \alpha K},\] and plugging this back into Equation 8.24 yields \[\E[S_2]=\E[N] \E[Z_1]=\frac{1}{2} \left( K+\frac{2}{\alpha} \right) e^{-\alpha K}\] and \[\begin{align*} \var(S_1) &= \var(N) \big( \E[Z_1] \big)^2 + \E[N] \var(Z_1) \\ &= \frac{6}{16} \left( K+\frac{2}{\alpha} \right)^2 e^{-2 \alpha K} + \frac{1}{2} \left( \frac{2K}{\alpha}+\frac{6}{\alpha^2} \right) e^{-\alpha K} - \frac{1}{2} \left( K+\frac{2}{\alpha} \right)^2 e^{-2 \alpha K} \\ &= \frac{1}{2} \left( \frac{2K}{\alpha}+\frac{6}{\alpha^2} \right) e^{-\alpha K} - \frac{1}{8} \left( K+\frac{2}{\alpha} \right)^2 e^{-2 \alpha K}. \end{align*} \] And ultimately (pffff) the premium we wanted is given by \[\pi(S_2)=\frac{1}{2} \left( K+\frac{2}{\alpha} \right) e^{-\alpha K} + \frac{\beta}{2} \left( \frac{2K}{\alpha}+\frac{6}{\alpha^2} \right) e^{-\alpha K} - \frac{\beta}{8} \left( K+\frac{2}{\alpha} \right)^2 e^{-2 \alpha K}.\]
Exercise 8.3 [***] We made in Example 8.2, where the number of risks that materialises was modelled by a Poisson distribution, the following spectacular observation: after applying an excess of loss risk sharing agreement, the number of non-zero payments that the excess provider Bark has to do also follows a Poisson distribution.
It turns out that this in fact is also true in a compound Binomial model! Indeed consider a collection of independent and identical risks, each of which may or may not materialise. Let \(N \sim \text{Binomial}(n,p)\) for some \(n \in \{1,2,\ldots\}\) and \(p \in (0,1)\) model the number of risks that materialise. Further let \(X_1, X_2, \ldots\) be an i.i.d. sequence of random variables modelling the losses suffered due to the risks that end up materialising. Suppose that an excess of loss risk sharing agreement is in place, with retention level \(K>0\) satisfying \(\P(X_i>K)>0\).
Let \(\widehat{N}\) denote the number of (non-zero) payments that the excess provider has to do. Show that \(\widehat{N}\) also has a Binomial distribution. What are its parameters?
One simple hint: follow exactly the same strategy as in the relevant part of Example 8.2!
We follow the same strategy as in Example 8.2. Recall from Section 7.2.2 e.g. that if a risk materialises with loss \(X_i\), then the excess provider needs to make a non-zero payment if and only if \(X_i>K\). So if \(N\) risks materialise, with losses \(X_1, \ldots, X_N\), then \(\widehat{N}\) equals the number of these that take a value larger than \(K\). And following the lead of Proposition 8.3, we can express this mathematically as \[\widehat{N}=\sum_{i=1}^N I_i, \tag{8.25}\] where \[I_i=\begin{cases} 1 & \text{if } X_i>K \\ 0 & \text{if } X_i \leq K. \end{cases} \]
In order to study \(\widehat{N}\), we use the observation from Proposition 8.3 that Equation 8.25 shows that \(\widehat{N}\) has a compound distribution, and that hence the tools from Section 7.3.2 can be used. In particular, from Theorem 7.1 we get that the moment generating function (mgf) of \(\widehat{N}\) is given by \[M_{\widehat{N}}(t) = M_N \big( \log(M_I(t)) \big), \tag{8.26}\] where \(M_N\) is the mgf of \(N \sim \text{Binomial}(n,p)\) and hence, from Definition 7.2, equals \[M_N(t)=\left( 1-p+pe^t\right)^n \quad \text{for all } t \in \R. \tag{8.27}\] Further we can work out the common mgf of the \(I_i\)’s by appealing to its def (cf. Section A.2) and working it out using the standard formula for the expectation of a function of a discrete random variable (cf. Equation A.9, with the choice \(h(x)=e^{tx}\)) — after all \(I_1\) is a discrete random variable with range/set of possible values \(R := \{0,1\}\): \[\begin{align*} M_I(t) &= \E[e^{t I_1}] \\ &= \E[h(I_1)] \\ &= \sum_{k \in R} h(k) \P(I_1=k) \\ &= h(0) \cdot \P(I_1=0) + h(1) \cdot \P(I_1=1) \\ &= 1 \cdot \P(X_1 \leq K)+e^t \cdot \P(X_1>K) \\ &= q e^t +1-q, \end{align*} \] where we defined \(q:=\P(X_1>K)\) (recall that \(\P(X_1>K)>0\) by assumption).
If we now plug this together with Equation 8.27 back into Equation 8.26, then we find that \[\begin{align*} M_{\widehat{N}}(t) &= M_N \big( \log(M_I(t)) \big) \\ &= \left( 1-p+pe^{\log(M_I(t))} \right)^n \\ &= \left( 1-p+p M_I(t) \right)^n \\ &= \left( 1-p+p (q e^t +1-q) \right)^n \\ &= \left( pq e^t +1-pq \right)^n. \end{align*} \] And there we are! :). If we compare this mgf with the form given in Definition 7.2, we see that it matches the mgf of a \(\text{Binomial}(n,pq)\) distribution. And since mgf’s uniquely determine distributions (cf. Section A.2), it follows that indeed \[\widehat{N} \sim \text{Binomial}(n,pq).\]