\[ \renewcommand{\P}{\mathop{\mathbb{P}}\nolimits} \newcommand{\E}{\mathop{\mathbb{E}}\nolimits} \newcommand{\var}{\mathop{\rm Var}\nolimits} \newcommand{\VaR}{\mathop{\rm VaR}\nolimits} \newcommand{\cte}{\mathop{\rm CTE}\nolimits} \newcommand{\cov}{\mathop{\rm Cov}\nolimits} \newcommand{\limsup}{\mathop{\rm limsup}} \newcommand{\liminf}{\mathop{\rm liminf}} \newcommand{\R}{\mathbb{R}} \newcommand{\Q}{\mathbb{Q}} \newcommand{\Z}{\mathbb{Z}} \newcommand{\N}{\mathbb{N}} \newcommand{\C}{\mathbb{C}} \renewcommand{\d}{\, \mathrm{d}} \newcommand{\dP}{\, \mathrm{d}\mathbb{P}} \newcommand{\eps}{\varepsilon} \renewcommand{\emptyset}{\varnothing} \]
6 Premium principles & risk measures
6.1 Why?
In Chapter 1–Chapter 5 we’ve mainly focussed on some important tools to deal with risks but haven’t yet done too much with the risks themselves. Time to start digging into that!
Commonly in risk management you have a particular reason to want to represent a risk i.e. future random loss represented by a non-negative random variable \(X\), in terms of a real number. Most commonly this happens in one of two possible contexts.
- Another party owns the risk \(X\) i.e. they will have to pay it when the time comes. They want to get rid of it and ask you to take over ownership of the risk, in exchange for a single, deterministic payment right now. So if you were to accept this deal, then it means for you that you gain an amount of money right now, called the premium1, but become liable to pay the random risk/loss \(X\) in the future (i.e. whatever value it ends up taking). Natural question: how large should the premium be for you to be willing to accept this deal? Some examples:
- If you are working for an insurance company and you sell say a car insurance product, then in essence this transaction means that the buyer of that product pays you a certain amount of money (the premium) and in exchange you promise to reimburse any damage to their car for say a year (the random risk/loss).
- In finance, a swap is a type of investment product that comes in several different forms but they all share the common principle that a buyer and seller agree to exchange between them a cashflow subject to uncertainty (the random risk/loss) and a certain cashflow (which is equivalent to a single fixed payment now i.e. a premium).
- A government is inviting building companies to tender2 for some large building project. Such projects are always to a smaller or greater extend subject to uncertainty, and hence so is how much it will cost a company to carry out the project (the random risk/loss). The company is wondering what the tender price (the premium) should be.
- You own the risk/loss \(X\) and you want to assess how much money you should set aside now and keep in reserve to be confident enough that when the risk materialises (i.e. \(X\) takes its value) you do not get into any financial trouble. This is for instance a very prominent question for any company whose activities are subject to some degree of uncertainty. If such a company wants to produces a financial report with an outlook for the coming year say, be it for internal use to decide on company strategy e.g. or for external publication, then such uncertain risks/losses need to be properly quantified in the form of such a reserve. Such a computed reserve can of course also more generally serve as a quantification of “how risky” \(X\) exactly is.
1 This term originates from insurance, but has been adopted in many other contexts as well
2 I.e. companies are invited to submit a plan for how they would make the project happen, and in particular the tender price: the amount of money they require to accept the project. The government then compares the plans that have been submitted and chooses which company to offer the project to
In this chapter we’ll discuss some general rules and principles that are often useful in addressing such questions. There is not a single “best” rule/principle, rather there are several interesting and prominent ones with their own strengths and weaknesses. Further our discussion here provides above all a good foundation to start thinking from rather than some instructive manual — in any particular application in practice you typically also want to take the specific circumstances on the ground on board to decide on the best way forward.
We discuss both contexts mentioned above together in one chapter because mathematically, they boil down to wanting to do the same thing: mapping risks/loss i.e. non-negative random variables to non-negative real numbers (namely, the premium in context 1 and the reserve in context 2) i.e. a mapping of the form \[\{ \text{Non-neg. random variables } X \} \to [0,\infty). \tag{6.1}\] Nevertheless, which of the two contexts we’re in is still crucially important: it dictates which mappings of this form are useful/sensible and also what properties you’d want to see. That’s why we discuss contexts 1 and 2 above in different sections.
Finally, a very important point to keep in mind: a random variable is a much more complicated object than a real number. That means that by applying a mapping, say \(m\), as in Equation 6.1, much information is lost: given that we know the number \(m(X) \in \R\) it is impossible to find out what \(X\) was. Indeed this is due to the fact that such \(m\) is necessarily many-to-one i.e. not injective: there are multiple (typically infinitely many) different non-negative random variables that all get mapped to the same number in \([0,\infty)\). This may ring a bell — indeed we saw exactly the same phenomenom, albeit in a completely different context, when we discussed “measures of dependence” in the previous chapter, cf. Section 5.1!
For convenience, throughout this chapter we use the following notation:
Definition 6.1 We define \[\mathcal{D}=\{ \text{Non-negative random variables} \}.\]
6.2 Premium principles
We start by focussing on context 1 mentioned in Section 6.1 above: what is a reasonable premium (price, fixed amount), you would demand in return for accepting a risk/loss modelled by a non-negative random variable \(X\) i.e. for accepting the obligation to pay whatever value \(X\) ends up taking when the time comes?
Remark 6.1. In a business/insurance context, the premium we’re talking about here is only that what is deemed necessary to cover \(X\). This is typically more precisely called the net premium. A company couldn’t survive on that amount alone: they also have cost of business to deal with such as a nice building for you to sit in in your fancy suit, they will want to put in a safety/profit margin as well (if market competition allows) etc. So the price that is ultimately quoted, called the gross premium, will be larger than the net premium. However we focus on net premiums only and just call it “premium”.
The rules/principles we’ll be discussing to address this question are mappings of the type Equation 6.1, but we give them a specific name to capture the context:
Definition 6.2 A premium principle is a rule, say \(\pi\), for assigning a premium (non-negative real number) to a risk/loss (non-negative random variable) i.e. a mapping of the type Equation 6.1: \[\pi: \mathcal{D} \to [0,\infty)\] (recall the notation from Definition 6.1).
Keep in mind the position you’re putting yourself in as “risk buyer”: you receive the deterministic/fixed amount \(\pi(X) \in [0,\infty)\), but in exchange you’ll have to pay the value that \(X\) ends up taking at some later point. Normally \(\pi(X)\) will be smaller than the largest value that \(X\) can take (the reasons for this are discussed under the “No rip-off” property in Section 6.2.2), so while you may very well end up earning money (if the value that \(X\) takes is less than \(\pi(X)\)) you may also very well end up losing money (if the value that \(X\) takes is larger than \(\pi(X)\))! It’s risky business, being a risk manager! ;).
6.2.1 Some prominent examples of premium principles
Here are some prominent examples of premium principles as defined in Definition 6.2 (though in some cases we may end up with \(\pi(X)=\infty\) for some \(X \in \mathcal{D}\)).
- Expected value principle: \[\pi(X)=(1+\beta) \E[X], \tag{6.2}\] where \(\beta \geq 0\) is called the loading factor.
- Variance principle: \[\pi(X)=\E[X]+\beta \var(X), \tag{6.3}\] where \(\beta \geq 0\) is (again) called the loading factor.
- Standard deviation principle: \[\pi(X)=\E[X]+\beta \sqrt{\var(X)}, \tag{6.4}\] where \(\beta \geq 0\) is (still!) called the loading factor.
- Exponential principle: \[\pi(X)=\frac{1}{\beta} \log \left( \E[e^{\beta X}] \right), \tag{6.5}\] where \(\beta > 0\) is called the risk aversion parameter.
- Esscher principle: \[\pi(X)=\frac{\E[X e^{\beta X}]}{\E[e^{\beta X}]}, \tag{6.6}\] where \(\beta > 0\) is a parameter (poor guy doesn’t even have a nice name!).
Note that 1, 2 and 3 are all of essentially the same nature: you take \(\E[X]\) as “baseline” and then you top that it up by a bit, where each principle does that topping up in a slightly different way. We’ll discuss the reason for taking \(\E[X]\) as “baseline” in Section 6.2.2 below (see under “Positive loading”). Principles 2 and 3 in particular add a top-up in terms of the variance of \(X\), which has a clear rationale: if \(X\) has a large variance, then its possible values are relatively widely spread out around its mean \(\E[X]\). So in particular for you as “risk buyer” there is the scary possibility that \(X\) takes a value much larger than its mean, which drives you to demand extra top-up i.e. a larger premium \(\pi(X)\).
The exponential principle comes from a Utility Theory background. If you’re familiar with this field you’ll remember that it assumes that every agent has a utility function expressing their individual risk appetite, and that they are willing to engage in a financial transaction if and only if increases (or at least doesn’t decrease) their level of (expected) utility. If such an agent uses the exponential utility function and you work out what premium they would minimally require to be willing to accept the risk/loss \(X\), you arrive at Equation 6.5.
The Esscher principle comes from the famous Esscher transform in Prob Theory: it transforms the original probability measure \(\P\) into a new one, say \(\widetilde{\P}\), which has in particular the property that it makes larger values of \(X\) more likely (and hence smaller values less likely). So \(\widetilde{\P}\) represents a more pessimistic/prudent worldview if you will, in which \(X\) is “more risky” than it was under \(\P\). If you then apply the Expected value principle from 1 above, with \(\beta=0\), i.e. you set \(\pi(X)=\widetilde{\E}[X]\) where \(\widetilde{\E}\) denotes expectation under \(\widetilde{\P}\), and you do the maths to translate \(\pi(X)\) back in terms of \(\P\)/\(\E\) then you end up with Equation 6.6.
Yeah, ok, but which one of these principles should I then actually use Kees?? As mentioned in Section 6.1, there is no good general answer to this question — it very much depends on the exact circumstances and is ultimately also always to a degree of making a(n experience led) choice. We’re “just” exploring sensible possibilities here!
Example 6.1 Let’s do a quick and simple explicit example. Take \(X \sim \text{Unif}(a,b)\) for some \(0 \leq a<b\). Recall from Appendix B that \[\E[X]=(a+b)/2 \quad \text{and} \quad \var(X)=(b-a)^2/12.\] Further we can compute for any \(\beta \in \R\), using Equation A.16 (and the pdf from Appendix B if you don’t remember it): \[\begin{align*} \E[e^{\beta X}] &= \int_a^b e^{\beta x} \frac{1}{b-a} \d x \\ &= \frac{e^{\beta b}-e^{\beta a}}{\beta (b-a)} \end{align*} \] and \[\begin{align*} \E[X e^{\beta X}] &= \int_a^b x e^{\beta x} \frac{1}{b-a} \d x \\ &= \frac{e^{\beta b} (b-1/\beta)-e^{\beta a} (a-1/\beta)}{\beta (b-a)}. \end{align*} \] If you’re not sure how to do the final integral above, just apply integration by parts i.e. \[\int_a^b u(x)v'(x) \d x = u(x)v(x) \big|_{x=a}^{x=b}-\int_a^b u'(x)v(x) \d x\] with \(u(x)=x\) and \(v'(x)=e^{\beta x}\).
Plugging this all into the above premium principles, we get the following values:
- \[\pi(X)=(1+\beta) \frac{a+b}{2},\]
- \[\pi(X)=\frac{a+b}{2}+\beta \frac{(b-a)^2}{12},\]
- \[\pi(X)=\frac{a+b}{2}+\beta \frac{b-a}{\sqrt{12}},\]
- \[\pi(X)=\frac{1}{\beta} \log \left( \frac{e^{\beta b}-e^{\beta a}}{\beta (b-a)} \right),\]
- \[\pi(X)=\frac{e^{\beta b} (b-1/\beta)-e^{\beta a} (a-1/\beta)}{e^{\beta b}-e^{\beta a}}.\]
6.2.2 Some desirable properties of premium principles
Now let’s move on to formulating some properties that premium principles ideally should have i.e. what properties would make them “reasonable”. Recall that we’re still in context 1 from Section 6.1: a premium principle \(\pi\) gives a premium (price, fixed amount) \(\pi(X) \in [0,\infty)\) for the obligation of having to pay the uncertain amount \(X \in \mathcal{D}\) later on.
Positive loading: \[\pi(X) \geq \E[X] \quad \text{for all } X \in \mathcal{D}.\] Intuitively it makes sense that you want as premium at least the expected amount you would be losing due to accepting \(X\) no? There is a bit more of a story here if you look at the bigger picture. Take a company which accepts many, say \(n\), such risks (independent copies of \(X\), say \(X_1,\ldots,X_n\)), like an insurance company selling many of the same insurance products to different people. Then \(\E[X]\) is an important threshold: if you were to set \(\pi(X)<\E[X]\), then for large \(n\) the probability that you end up losing money i.e. that your total outgo \(X_1+\ldots+X_n\) turns out to be larger than your total premium income \(n \pi(X)\) is quite large, increasing to \(1\) even if \(n \to \infty\)! On the other hand, if you set \(\pi(X)>\E[X]\) then that same probability is quite small and decreasing to \(0\) if \(n \to \infty\)! See also Exercise 6.5.
Additivity: \[\pi(X+Y)=\pi(X)+\pi(Y) \quad \text{for all \emph{independent} } X,Y \in \mathcal{D}.\] Logically, if a risk is the sum of two independent components then the premium for them together should simply be the sum of premiums of the components no.
Consistency: \[\pi(X+c)=\pi(X)+c \quad \text{for all } X \in \mathcal{D} \text{ and constants } c \geq 0.\] If a risk/future payment consists of the sum of a fixed amount \(c\) plus a random amount \(X\), then logically the premium should be the sum of \(c\) (covering the \(c\) you’ll know you’ll have to pay, making this part a straightforward swap of money) and \(\pi(X)\) no.
No rip-off: for any \(X \in \mathcal{D}\), if a constant \(C\) exists so that \(X \leq C\) then \(\pi(X) \leq C\).
If a risk has an upper bound \(C\) i.e. you know that its value will never be more than \(C\), then it doesn’t make a lot of sense to charge a premium larger than \(C\) no? I mean, you could try, but anybody willing to enter that deal with you should probably better hand in their wallet with their mum — after all they would obviously be better off just paying \(X\) themselves when the time comes rather than paying you \(\pi(X)>C \geq X\) to get rid of the risk!3
Scale invariance: \[\pi(cX)=c \pi(X) \quad \text{for all } X \in \mathcal{D} \text{ and constants } c > 0.\] This property expresses that premiums should scale in the obvious way with the risks. On the one hand it’s one that a slightly autistic mathematical mind finds very appealing and neat, but on the other hand, of the properties we’re discussing here this is probably the one that is most debatable. Larger risks are extra scary if you will because of the potential to cause serious financial damage and that translates in practice often to a larger premium than a proportionally scaled one.
Monotonicity: \[\pi(X) \leq \pi(Y) \quad \text{for all } X,Y \in \mathcal{D} \text{ satisfying } X \leq Y.\] If we know that the value that \(Y\) takes is guaranteed to be at least as large as the value that \(X\) takes, then it only makes sense that the premium for \(Y\) is at least as large as the premium for \(X\) right.
3 There’s an argument from Economy/Mathematical Finance as well: if it were true that \(\pi(X)>C\) then the risk buyer would make a guaranteed profit, which is in Mathematical Finance called an arbitrage opportunity. Economical theory dictates that under the assumption of a transparent and open market filled by rational agents such opportunities cannot exist, or at least not for very long: it would quickly catch the eye of other agents who also would want to set up such deals, creating an increase in demand for the risk \(X\), which by the laws of supply and demand would increase the price of the risk \(X\) until there is no more free money to be made
That’s our list! You’ll hopefully agree with me that these are all sensible properties to want your premium principle to have. And that at least the premium principles 1, 2 and 3 from Section 6.2.1 are sensible examples of premium principles (4 and 5 are naturally harder to judge if you’re not familiar with their origins). However, it turns out that it the world is not always as well behaved as we would like it to be — even the sensibly looking premium principles do not necessarily satisfy every sensible property!
Example 6.2 Let’s take the variance principle, i.e. principle 2 from Section 6.2.1: \[\pi(X)=\E[X]+\beta \var(X) \quad \text{for all } X \in \mathcal{D},\] where we fix a loading factor \(\beta>0\). Let’s walk through the above six properties and see whether or not they hold for this \(\pi\)!
- Holds obviously.
- Holds as well, after all \(\E[X+Y]=\E[X]+\E[Y]\) always holds, and we have that \(\var(X+Y)=\var(X)+\var(Y)\) as well by independence of \(X\) and \(Y\) (recall Equation A.55).
- Also holds! Just recall that if \(c\) is a constant, then \(\var(X+c)=\var(X)\). You’ll probably remember this. Intuiitvely because the variance measures spread, and the spread doesn’t change if we move all values that our random variable can take up by the same amount \(c\). Mathematically, just work out the def of variance.
So far so good! But it starts going downhill from here sadly. How do we show/prove that a premium principle \(\pi\) does not satisfy a certain property? Well since these properties are of the form “for all [blah], \(\pi\) does [blahblah]”, to prove that this doesn’t hold a general strategy is to prove the logical negation: “there exists a [blah] for which \(\pi\) does not do [blahblah]”. Typically this boils down to finding an example of an \(X \in \mathcal{D}\) so that \(\pi(X)\) does not satisfy a certain property. General advice here is to just try the simplest \(X \in \mathcal{D}\) you can think of, like a random variable taking only two possible values — keeps your life as easy as it gets! If that doesn’t work out, you can always sulk and move on to other more complicated examples.
We claim that this one does not hold. To prove this, we need to find an \(X \in \mathcal{D}\) and a constant \(C\) satisfying \(X \leq C\), for which \(\pi(X) >C\). As advised above, let’s take the simplest possible (non-trivial) example of \(X \in \mathcal{D}\), namely \[X=\begin{cases} 0 & \text{with prob } 1/2 \\ a & \text{with prob } 1/2 \end{cases} \] for some \(a>0\). We leave \(a\) variable for now so that later we still have the freedom to choose its value and get the result we want. If you don’t think you need that, you could of course just take \(a=1\) e.g. On the other hand, if you think/find that you need more freedom, you could for instance leave the probabilities variable (i.e. \(p\) and \(1-p\) for \(p \in (0,1)\)).
Ok, let’s carry on. We can choose \(C=a\), then \(X \leq C\) is clearly satisfied. Further we can easily compute (recall Equation A.10 and Equation A.9) \[\E[X]=0 \cdot \frac{1}{2} + a \cdot \frac{1}{2}=\frac{a}{2}\] and \[\E[X^2]=0^2 \cdot \frac{1}{2} + a^2 \cdot \frac{1}{2}=\frac{a^2}{2}\] so that \[\var(X)=\E[X^2]-\big( \E[X] \big)^2=\frac{a^2}{4}\] and hence \[\pi(X)=\frac{a}{2}+\beta \frac{a^2}{4}.\] So in order to wrap up our argument here, we are looking to find a value for \(a>0\) so that \(\pi(X)>C=a\) i.e. so that \[\frac{a}{2}+\beta \frac{a^2}{4}>a \iff \frac{a}{4}(\beta a-2)>0.\] So any choice \(a>2/\beta\) works!
More bad news — also doesn’t hold! You could very well follow the same strategy as in 4 above, but we can also see this more generally. Indeed for any \(c>0\) we can work out \[\begin{align*} \pi(cX) &= \E[cX]+\beta \var(cX) \\ &= c \E[X]+ \beta c^2 \var(X) \end{align*} \] while \(c\pi(X)=c\E[X]+c \beta \var(X)\), so indeed for any \(c \not= 1\) these are not equal!
Guess what — also not true! Along the same lines as in 4, we show this by finding examples of random variables \(X,Y \in \mathcal{D}\) so that \(X \leq Y\) but \(\pi(X)>\pi(Y)\). Again we look for the simplest possible examples first: let’s use the same \(X\) as in 4 i.e. \[X=\begin{cases} 0 & \text{with prob } 1/2 \\ a & \text{with prob } 1/2 \end{cases} \] and for \(Y\) we simply choose the trivial constant random variable \(Y=a\). (Even simpler would be to choose both random variables constant, but that doesn’t yield the result we’re looking for.) Then indeed \(X \leq Y\) holds, and it remains to try and find a choice for \(a>0\) so that \(\pi(X)>\pi(Y)\) holds. We had already computed \[\pi(X)=\frac{a}{2}+\beta \frac{a^2}{4}\] and since \(\E[Y]=a\) and \(\var(Y)=0\) (if you’re not sure, just work out the def) \[\pi(Y)=a.\] So we’re looking for an \(a>0\) so that \[\frac{a}{2}+\beta \frac{a^2}{4}>a.\] Coincidentally this is the same inequality we were staring at in 4 above, and we saw there already that any \(a>2/\beta\) does the trick!
6.3 Risk measures
Now we move on to context 2 mentioned in Section 6.1: we’re facing a certain risk/loss \(X\) (a non-negative random variable) in the future, and we’re wondering how much money we should set aside i.e. keep in reserve to be confident enough that we will be able to deal with this loss. As already mentioned: this can apply literally, but the required reserve that is computed can also act as a measure for how “risky” \(X\) is rather (hence the term risk measure!). As was the case for premium principles, there is a simple general definition:
Definition 6.3 A risk measure is a rule, say \(\rho\), for assigning a required reserve (non-negative real number) to a risk/loss (non-negative random variable) i.e. a mapping of the type Equation 6.1: \[\rho: \mathcal{D} \to [0,\infty)\] (recall the notation from Definition 6.1).
As mentioned before: though mathematically the same object (just compare Definition 6.2 with Definition 6.3), the context and interpretation are different than the viewpoint we took in Section 6.2. This also means that for this purpose, we tend to use different mappings \(\mathcal{D} \to [0,\infty)\) than we did in Section 6.2 etc.
Here’s one obvious question (we asked the analogue question for premium principles in terms of the “No Rip-off” property discussed in Section 6.2.2): if you’re so worried about that risk \(X\), why then not simply keep enough money in reserve so that you’re always safe i.e. set it to the largest value that \(X\) can possible take and be done with it? Sure, that’s (normally) an option! However this also has its downsides. For example, any money kept in reserve cannot be used for other purposes. It cannot be externally invested to make an attractive investment return (at least not in any way that risks loss in value, leaving only extremely safe and therefore normally minimally performant investment opportunities). It cannot be internally invested in research & development, or used for initiatives to grow the business etc. or any other purpose that is important for the future of the business. It can even be a counterproductive strategy, by cutting off ways to generate extra revenue/profit and thereby possibly provide a better guarantee that there is enough money available to pay \(X\) than just putting it in a safe somewhere!
For such reasons it is worthwhile to consider keeping less in reserve than the largest possible value of \(X\). But, of course, that does inevitably come with the risk that \(X\) takes a value larger than the amount you’ve kept in reserve!
In this section, we will introduce and briefly the discuss the most prominent risk measures. Because of their wide use (e.g. in financial reporting) you may well have encountered them before, also outside your studies!
Remark 6.2. In principle there is no reason for the domain of a risk measure to be restricted to non-negative random variables/risks only, there is no objection to apply them to any random variable \(X\) in principle. However the perspective we always take is that \(X\) is a risk/loss, so if \(X\) takes a positive value then that means that the “risk holder” looses money while if it takes a negative value (in cases where we allow that) it means that they gain money.
6.3.1 Value-at-Risk (VaR)
The Value-at-Risk (VaR) is the simplest and also most prominent risk measure. You specify a (confidence) level \(p \in (0,1)\) (normally close to \(1\)), and then you’re looking for the (smallest) cut-off point in the range of \(X \in \mathcal{D}\) so that the probability that \(X\) takes a value larger than this cut-off point is (at most) \(1-p\). This cut-off point is then your VaR.
Now, this talk of “cut-off point” may ring a bell with you — indeed we looked at them before, albeit in a different context, in Section 5.4.1! In particular, using Lemma 5.1 we can immediately write down an expression for the cut-off point described in the previous paragraph, namely \(F_X^{-1}(p)\).
Definition 6.4 Let \(X \in \mathcal{D}\) (cf. Definition 6.1) with cdf \(F_X\) and \(p \in (0,1)\) (normally close to \(1\)). Then the Value-at-Risk for \(X\) at (confidence) level \(p\) is defined as \[\VaR_p(X)=F_X^{-1}(p) \in [0,\infty)\] (where, I’m sorry, we may have to use the generalised inverse cdf again if needed, recall from Section 2.2.1).
Its meaning: if we decide to keep the amount \(\VaR_p(X)\) in reserve, then we have enough to cover the future value of \(X\) with probability (at least) \(p\).
Note:
The use of VaR is very wide spread: banks, insurers and other financial companies use them extensively to assess their risks, both internally and externally. Large regulatory frameworks such as
- Solvency II (for insurers, EU wide),
- Solvency UK (for insurers, the UK’s somewhat adjusted Solvency II since Brexit),
- C-ROSS (for insurers in China),
- Basel II (for banks and related, pretty much worldwide)
use (or used) them to prescribe how much risk companies covered by such frameworks are allowed to take.
You can maybe already hear the “but” coming ;). Fundamentally, any risk measure will always have limitations — that’s an inevitable consequence of the fact that they can only represent the risk/loss to a limited extent (recall the “loss of information” mentioned in Section 6.1). A particular limitation of the VaR is that it only gives you that cut-off point. It doesn’t provide any information about how bad things can get in the unlikely (for \(p\) close to \(1\)) yet not impossible event that \(X\) does take a value larger than the VaR.
A number of smaller but also quite prominent financial disasters, such as
- in 2000, the UK insurance company Equitable Life which had as many as 1.5 million policyholders and managed funds to the tune of £26 billion came very close to collapsing, leading to the Morris Review;
- in 2008, systemic problems in the US housing market envolved into a full blown global financial crisis that lasted several years, brought down a number of prominent companies in the financial industry (such as the investment back Lehman Brothers in the US) and required governments worldwide to bail out many others that were on the brink of collapse;
resulted in a series of revisions of existing laws, rules, best practices and regulatory frameworks. Though such events hardly ever have only one single cause of course, the over-reliance on VaR based thinking in risk management and in particular not enough consideration for its limitations has in many investigations been identified as an important factor.
Inspired by this, alternatives to the VaR have gained a lot of traction in recent decades. For example, the Basel II framework mentioned above was replaced by Basel III, finalised in 2010, in which risk measures such as the Conditional Tail Expectation (and similar) play a much more important role.
6.3.2 Conditional Tail Expectation (CTE) and others
We mentioned in the above section already that the VaR gives a cut-off point that \(X\) exceeds with a small probability only, which is useful, but it doesn’t give any information about how large the values of \(X\) can be if we do happen to end up above that cut-off point. The Conditional Tail Expectation addresses that, by computing the expectation/average over those values of \(X\) which exceed that cut-off point:
Definition 6.5 Let \(X \in \mathcal{D}\) (cf. Definition 6.1) with \(\P(X>\VaR_p(X))>0\) and let \(p \in (0,1)\) (normally close to \(1\)). The Conditional Tail Expectation (CTE) at level \(p\) is given by \[\cte_p(X)=\E \big[X \, \big| \, X>\VaR_p(X) \big] \in [0,\infty],\] where \(\VaR_p(X)\) is the Value-at-Risk as defined in Definition 6.4.
Some notes:
- We mean by \([0,\infty]\) simply \([0,\infty) \cup \{\infty\}\), so any non-negative real number plus \(\infty\). For an example where this \(\infty\) actually pops up, see the Pareto distribution in Example 6.4 and Exercise 6.4.
- Recall from Definition 6.4 that if \(X\) is continuous, then \(\P(X>\VaR_p(X))=1-p>0\) and hence this definition can be used.
- If \(\P(X>\VaR_p(X))=0\) then \(\cte_p(X)\) is not well-defined. Intuitively: if \(X\) does not take any values larger than the cut-off point \(\VaR_p(X)\), then we definitely cannot compute the expectation over these non-existing values to form the CTE! A trivial example of this is when \(X\) is constant.
Example 6.3 Here is quick and simple example so that you can hopefully clearly follow what happens. Let \(X\) be a random variable with range \(\{1,\ldots,6\}\) and \(\P(X=k)=1/6\) for \(k=1,\ldots,6\) (rolling a die eh!). Fix (say) \(p=0.6\). Let’s find VaR and CTE using the descriptions we’ve given rather than using their mathematical expressions from Definition 6.4 and Definition 6.5.
For \(\VaR_p(X)\), recalling the description above Definition 6.4: we’re looking for the (smallest) cut-off point in the range of \(X\) i.e. \(\{1,\ldots,6\}\) so that the probability that \(X\) takes a value larger than this cut-off point is (at most) \(1-p\). We claim that for our choice \(p=0.6\) this cut-off point is \(4\). Indeed, \[\P(X>4)=\frac{1}{3} \leq 1-p=0.4\] while \[\P(X>3)=\frac{1}{2} > 1-p=0.4.\] So \[\VaR_{0.6}(X)=4.\]
For \(\cte_{0.6}(X)\), as mentioned above Definition 6.5, we’re looking for the expectation/average over those values of \(X\) which exceed that cut-off point \(\VaR_{0.6}(X)=4\). Clearly these are the values \(5\) and \(6\), which \(X\) takes with equal probability, so that the CTE is their (simple) average: \[\cte_{0.6}(X)=\frac{5+6}{2}=5.5.\]
If \(X\) represented the outcome of an unfair die, say that it takes value \(5\) with prob \(q\) and value \(6\) with probability \(1/3-q\) for some \(q \in (0,1/3)\) while the other probs remain unchanged (so that the VaR value doesn’t change), to compute CTE we need to account for these different probabilities for \(5\) and \(6\) by taking the properly weighted average \[\begin{align*} \cte_{0.6}(X) &= \frac{q}{q+\frac{1}{3}-q} \cdot 5 + \frac{\frac{1}{3}-q}{q+\frac{1}{3}-q} \cdot 6 \\ &= 3(2-q) \end{align*} \] (note that the denominator we used in the fractions ensures that we proportionally scale up the original weights/probs \(q\) and \(1/3-q\) so that the sum of the scaled weights is \(1\)). Indeed for \(q=1/6\) we find \(\cte_{0.6}(X)=3 \cdot (2-1/6)=5.5\) back!
Final note: if we were to change to (say) \(p=0.9\), then \(\VaR_{0.9}(X)=6\) (because \(\P(X>6)=0 \leq 1-p=0.1\) but \(\P(X>5)=1/3 > 1-p=0.1\)). So in this situation, \(\VaR_{0.9}(X)\) equals the largest possible value that \(X\) can take. This also means that it is impossible for \(X\) to take a value larger than its VaR and hence we can’t make any sense of what the CTE should be. Indeed now the condition needed in Definition 6.5, i.e. that \(\P \big( X>\VaR_{0.9}(X) \big)>0\), is not satisfied and hence the CTE is not defined (this is an example of what the final bullet point in Definition 6.5 is referring to).
Since the CTE is the (weighted) average/expectation of those values of \(X\) that exceed the cut-off point \(\VaR_p(X)\), it seems clear that the result should always be larger than \(\VaR_p(X)\):
Lemma 6.1 For any \(X \in \mathcal{D}\) with \(\P(X>\VaR_p(X))>0\) (recall the notes in Definition 6.5) and \(p \in (0,1)\) we have that \[\cte_p(X) > \VaR_p(X).\]
Proof. We have (don’t forget, \(\VaR_p(X)\) is just a real number/constant!): \[\begin{align*} \cte_p(X) &= \E \big[X \, \big| \, X>\VaR_p(X) \big] \\ &= \frac{\E \left[\mathbf{1}_{\{X>\VaR_p(X)\}} X \right]}{\P(X>\VaR_p(X))} \\ &> \frac{\E \left[\mathbf{1}_{\{X>\VaR_p(X)\}} \VaR_p(X) \right]}{\P(X>\VaR_p(X))} \\ &= \frac{\VaR_p(X) \E \left[\mathbf{1}_{\{X>\VaR_p(X)\}} \right]}{\P(X>\VaR_p(X))} \\ &= \frac{\VaR_p(X) \P(X>\VaR_p(X))}{\P(X>\VaR_p(X))} \\ &= \VaR_p(X), \end{align*} \] where the second line uses Equation A.43 and the fifth line uses Equation A.11.
So it remains to justify the third line. Bit technical, sorry — don’t worry about this for the exam. Let’s write \(c=\VaR_p(X)\) for simplicity, so that we need to show that \[\E \left[\mathbf{1}_{\{X>c\}} X \right]>\E \left[\mathbf{1}_{\{X>c\}} c \right]\] i.e. that \[\E \left[\mathbf{1}_{\{X>c\}} (X-c) \right]>0. \tag{6.7}\] The weak inequality is quite easy to see, after all, recalling that random variables are mappings \(\Omega \to \R\) (recall from Section A.1), for any \(\omega \in \Omega\) we have that \(\mathbf{1}_{\{X(\omega)>c\}} (X(\omega)-c) \geq 0\) and hence Equation 6.7 holds with an “\(\geq\)” for sure. The argument for the strict inequality is a bit harder.
Maybe you have seen in an earlier Prob course the following result: if a non-negative random variable \(Y\) satisfies \(\P(Y>0)>0\), then \(\E[Y]>0\). If so, then you can apply this and be done with it. Otherwise, if you have only seen the result that if \(Y\) is a non-negative random variable then \(\E[Y] \geq 0\) (i.e. the argument we used above to argue that the weak inequality is easy to see), then there is some work to be done, for instance as below.
First we claim that there exists an \(n \in \N\) so that \(\P(X>c+1/n)>0\). Indeed, if this were not true, then \[\begin{align*} \P(X>c) &= 1-F_X(c) \\ &= 1- \lim_{n \to \infty} F_X \left( c+\frac{1}{n} \right) \\ &= \lim_{n \to \infty} 1- F_X \left( c+\frac{1}{n} \right) \\ &= \lim_{n \to \infty} \P \left( X>c+\frac{1}{n} \right) \\ &= \lim_{n \to \infty} 0 \\ &= 0, \end{align*} \] contradicting the \(\P(X>c)>0\) we assumed in the statement of the lemma. Note that in the second equality we used right-continuity of the cdf \(F_X\) (recall Section A.1.1).
Now we can fix an \(n \in \N\) with \(\P(X>c+1/n)>0\) and manipulate the left hand side of Equation 6.7 as follows: \[\begin{align*} \E \left[\mathbf{1}_{\{X>c\}} (X-c) \right] &= \E \left[\mathbf{1}_{\{X>c\}} (\mathbf{1}_{\{X \leq c+1/n \}}+\mathbf{1}_{\{X > c+1/n \}}) (X-c) \right] \\ &= \E \left[\mathbf{1}_{\{X>c\}} \mathbf{1}_{\{X \leq c+1/n \}} (X-c) \right] \\ & \qquad + \E \left[\mathbf{1}_{\{X>c\}} \mathbf{1}_{\{X > c+1/n \}} (X-c) \right] \\ &= \E \left[\mathbf{1}_{\{X \in (c,c+1/n]\}} (X-c) \right] \\ & \qquad + \E \left[\mathbf{1}_{\{X>c+1/n\}} (X-c) \right] \\ &\geq 0 + \E \left[\mathbf{1}_{\{X>c+1/n\}} \frac{1}{n} \right] \\ &= \frac{1}{n} \P(X>c+1/n)>0, \end{align*} \] where the inequality uses that in the first expectation, the random variable is non-negative and for the second expectation that \(\mathbf{1}_{\{X>c+1/n\}} (X-c) \geq \mathbf{1}_{\{X>c+1/n\}} 1/n\), and the final equality uses Equation A.11.
The CTE is not the only alternative(/improvement?) to the classic VaR, there is for instance also the Expected Shortfall (ES) and the Tail-Value-at-Risk (TVaR). These both work in a similar spirit as the CTE (and in fact, the TVaR is typically equal to the CTE). If you’re interested in more, see e.g. Section 2.2 in McNeil, Frey, and Embrechts (2015) or Chapter 5 in Kaas et al. (2008).
Remark 6.3. Besides the issues addressed above, the VaR has another downside: it is not subadditive. Any risk measure, say \(\rho\), is called subadditive if it holds that \[\rho(X+Y) \leq \rho(X)+\rho(Y) \quad \text{for all } X,Y \in \mathcal{D}. \tag{6.8}\] A simple example which confirms that VaR is not subadditive is for instance the following. Let \(X,Y\) be independent with the same distribution: they take value \(0\) with prob \(0.99\) and value \(100\) with prob \(0.01\). Then \(X+Y\) has possible values \(0, 100, 200\) with probs \[\P(X+Y=0)=0.99^2, \quad \P(X+Y=100)=2 \cdot 0.01 \cdot 0.99, \quad \P(X+Y=200)=0.01^2.\] Further (check for instance in the same way as in Example 6.3, or find the inverse cdf of \(X\) and \(Y\) of course) \[\VaR_{0.99}(X)=\VaR_{0.99}(Y)=0,\] but \[\VaR_{0.99}(X+Y)=100\] since \(\P(X+Y>100)=0.01^2 \leq 1-0.99\) while \(\P(X+Y>0)=1-0.99^2 > 1-0.99.\)
This is not just a mathematical inconvenience. Imagine that you and your friend work for the same company and that you are responsible for managing risk \(X\) and they for risk \(Y\). Your manager asks how much they should keep in reserve to cover your risks, you tell him \(\VaR_p(X)\) and your friend \(\VaR_p(Y)\). It doesn’t make a lot of sense if your manager, who is responsible for both risks, then actually needs to keep more in reserve than the sum \(\VaR_p(X)+\VaR_p(Y)\) no?
Another way to look at this is the concept of risk diversification: an important general wisdom in risk management is that it is best not to put all your eggs in one basket but to have smaller risks in different baskets rather. It reduces the risk of a catastrophic failure: if something goes horribly wrong with one of the baskets, then it is better to loose some of your eggs rather than all of them eh? A subadditive risk measure \(\rho\) expresses this: if you hold the risks \(X\) and \(Y\) both, then your risk level \(\rho(X+Y)\) is smaller (or at least not larger) than the sum of the risk levels.
The CTE on the other hand, is subadditive — a mathematical proof requires some work that we don’t want to get into here. If you’re interested in the maths involved (which is quite interesting I think!), see e.g. Embrechts and Wang (2015).
Let’s conclude with another nice example, this time considering several continuous distributions!
Example 6.4 First let’s take \(X \sim \text{Unif}(0,1)\), and fix some level \(p \in (0,1)\) (think: close to \(1\)). We already know what its inverse cdf is, namely \(F_X^{-1}(y)=y\) for all \(y \in (0,1)\) (recall from Example 3.4 e.g.) and hence we get immediately from Definition 6.4 that \[\VaR_p(X)=F_X^{-1}(p)=p. \tag{6.9}\] It’s also not very hard to work out the CTE from Definition 6.5: \[\begin{align*} \cte_p(X) &= \E \big[X \, \big| \, X>\VaR_p(X) \big] \\ &= \frac{\E \left[\mathbf{1}_{\{X>p\}} X \right]}{\P(X>\VaR_p(X))} \\ &= \frac{1}{1-p} \int_p^1 x \d x \\ &= \frac{1}{1-p} \frac{1-p^2}{2} \\ &= \frac{1+p}{2}, \end{align*} \] where for the second equality we used Equation A.43 and plugged in Equation 6.9, and for the third Equation A.45 with \(I=(p,\infty)\) (if you had forgotten the pdf of \(X\), keep Appendix B in mind) as well as that by very def \(\P(X>\VaR_p(X))=1-p\) (recall from Definition 6.4).
Sanity check (always recommended!):
- Recall that the VaR gives the cut-off point in the range of \(X\) (so in this case the interval \((0,1)\)) so that the probability that \(X\) is larger than this cut-off point is \(1-p\), that is: we’re looking for \(x \in (0,1)\) so that the probability that \(X\) ends up in the interval \((x,1)\) is \(1-p\). Due to the fact that all values in \((0,1)\) are “equally likely” that interval should have length \(1-p\) i.e. the cut-off point is \(x=p\).
- As discussed, intuitively the CTE gives us the expected value/weighted average of the values of \(X\) that are larger than the VaR i.e. \(p\) in this case. For this \(\text{Unif}(0,1)\) distribution, these values make up the interval \((p,1)\) and are all “equally likely”, therefore their expected value/weighted average is naturally right in the middle of that interval i.e. \((1+p)/2\).
Note also that \(\cte_p(X)>\VaR_p(X)\), as Lemma 6.1 told us should be the case, but they are not very far away from each other. In fact their difference vanishes as we let \(p \uparrow 1\). If you think about it, then this will always happen if we are dealing with a (continuous) distribution that has a bounded range i.e. if the possible values of \(X\) are bounded above.
Next let’s take \(Y \sim \text{Exp}(\theta)\) for some \(\theta>0\) and again fix some level \(p \in (0,1)\). Leaving the dirty work for you to do in Exercise 6.4, we can compute \[\VaR_p(Y)=\frac{-1}{\theta} \log(1-p). \tag{6.10}\] and \[\cte_p(Y)=\frac{\big( 1-\log(1-p) \big)}{\theta}. \tag{6.11}\] Since \(Y\) has range \((0,\infty)\) (if you’re not sure, recall Equation A.18), we would expect that the VaR at level \(p\), being the cut-off point so that the probability that \(X\) exceeds it is \(1-p\), tends to \(\infty\) as we let \(p \uparrow 1\). And that’s indeed the case as you can readily verify from Equation 6.10. Further note that in this case \[\cte_p(Y)=\frac{1}{\theta}+\VaR_p(Y),\] which not only confirms Lemma 6.1 but in particular shows that the distance between the VaR and the CTE is independent of \(p\): if we let \(p\) grow to \(1\) then both go to \(\infty\) (since the VaR does as we just discussed) but as they do so, they nicely keep a constant distance \(1/\theta\) between each other. Note how this contrasts with the \(\text{Unif}(0,1)\) case above, where the distance between VaR and CTE vanishes as \(p \uparrow 1\).
(Aside: if you have a very sharp eye, you may have noticed that this distance \(1/\theta\) is also exactly the mean of \(Y\). This is not a coincidence: we could have predicted this solely on the basis of the lack of memory property of the Exponential distribution! This also means that the Exponential distribution is the only continuous distribution where the distance between VaR and CTE is independent of \(p\)).
One more then! Let \(Z\) have a Pareto distribution with parameters \(x_m=1\) and \(\alpha>0\), which has pdf \[f_Z(x)=\begin{cases} \alpha x^{-\alpha-1} & \text{if } x>1 \\ 0 & \text{if } x \leq 1, \end{cases} \tag{6.12}\] with hence range \((1,\infty)\) (cf. Equation A.18 if you’re not sure). This is an example of a so-called power law: as \(x\) grows towards \(\infty\) (i.e. where the large and hence dangerous values of \(Z\) are!) then \(f_Z(x) \downarrow 0\), as all (reasonable/relevant for us) pdf’s do, but it does so only relatively slowly. In particular, compare this with for instance the pdf of an Exponential, Gamma or Normal distribution (cf. Appendix B): these all decay exponentially fast and that is fundamentally faster than a power function like \(f_Z\) does.
This makes such power laws basically a risk manager’s worst nightmare: roughly speaking, the probability that it takes a very large value (and hence results in a massive liability!) is fundamentally larger than any of those distributions with exponential decay manage to achieve. Or, a bit more precisely: for any \(p \in (0,1)\) close to \(1\), if \(Z\) takes a value exceeding the cut-off point \(\VaR_p(Z)\) then this tends to be a much larger value than a random variable with the same VaR value but with an exponentially decaying pdf would.
It was exactly warning us about this aspect that we introduced the CTE for, so let’s see if it shows up for the job! Again leaving the sweaty bit up to you in Exercise 6.4, we have that \[\VaR_p(Z)=(1-p)^{-1/\alpha} \tag{6.13}\] and \[ \cte_p(Z) = \begin{cases} \infty & \text{if } \alpha \leq 1 \\ \frac{\alpha}{\alpha-1} (1-p)^{-1/\alpha} & \text{if } \alpha>1. \end{cases} \tag{6.14}\] Wow — so for \(\alpha \in (0,1]\), for any \(p \in (0,1)\), the average/expected value over those values of \(Z\) that exceed \(\VaR_p(Z)\) is even infinite! Contrast this with the suddenly very pale and well behaved Exponential distribution \(Y\), where that same quantity is only a constant \(1/\theta\) larger than its VaR. But even if \(\alpha>1\), the difference \[\begin{align*} \cte_p(Z)-\VaR_p(Z) &= \frac{\alpha}{\alpha-1} (1-p)^{-1/\alpha}-(1-p)^{-1/\alpha} \\ &= \frac{1}{\alpha-1} (1-p)^{-1/\alpha} \end{align*} \] explodes to \(\infty\) as \(p \uparrow 1\). Again, contrast this with the constant difference for the Exponential distribution, and that difference going to \(0\) even in the case of a Uniform distribution (or any other distribution with an upper bound on its range)!
Finally, if you’d prefer a like-for-like numerical comparison, take for instance the following. Let \(p=0.95\), \(\theta=0.2\) as parameter for \(Y\), and \(\alpha=1.1\) as parameter for \(Z\). Then we get from Equation 6.10 and Equation 6.13 that \[\VaR_{0.95}(Y)=\frac{-1}{0.2} \log(0.05) \approx 15 \quad \text{and} \quad \VaR_{0.95}(Z)=0.05^{-1/1.1} \approx 15,\] so both \(Y\) and \(Z\) exceed \(15\) with (approx) equal probability, namely \(0.05\). Yet from Equation 6.11 and Equation 6.14 we get that \[\cte_{0.95}(Y)=\frac{1-\log(0.05)}{0.2} \approx 20 \quad \text{and} \quad \cte_{0.95}(Z)=\frac{1.1}{0.1} 0.05^{-1/1.1} \approx 168,\] so despite these (approx) identical probabilities, if we know that \(Z\) exceeds \(15\) then its expected value is (approx) \(168\) while for \(Y\) that is only (approx) \(20\)!
Another side of this same coin: if a VaR focused company has as rule that for a risk, they reserve an amount of four times the VaR (this factor of four is pretty arbitrary, but it was/is(?) not all uncommon to see companies do this stuff — “it feels nice and comfortably large, so it should surely be good enough!” style of thinking), i.e. (approx) \(60\) in the case of both \(Y\) and \(Z\), then the probability that it would nevertheless still end up short of money is \[\P(Y>60) \approx 0.000006 \quad \text{vs.} \quad \P(Z>60) \approx 0.01\] i.e. virtually impossible vs. a much more scary one in every hundred times…
6.3.3 Alternative risk management methods
We conclude with a quick wider comment. A practical difficulty with risk measure based risk management can be getting the mathematical modelling of your risks right (enough), in particular when it comes to extreme events: it’s notoriously hard to accurately estimate the probability of very unlikely events4 as well as it’s not easy to accurately understand just how wrong things can go if they really go wrong (think serious financial crisis style, with bankruptcies, governments/regulators being stretched in terms of what they can compensate/save etc.). Another possible issue is that, especially in economically unstable times, risk measures and hence levels of reserves required for complicated portfolios of assets and liabilities can easily change value quick and often. This hampers the fluency of doing business, predictability of capital available for investment, hurt outside perception of a company’s financial viability etc.
4 After all, estimation of probabilities is often based on empirical observations, and if you hardly ever get to observe something…
Such considerations have led, especially also in the aftermath of the 2008 financial crisis alluded to in Section 6.3.1, to discussions about moving away from strictly mathematical approaches to risk management in favour of leverage based approaches for instance. At the end of the day, no approach will ever be perfect and a responsible company/regulator/government will use a mix of methods implemented by a mix of talented young people (like you!) and experienced people (like you a while from now!) to inform and guide them in their risk management strategies.
6.4 Some exercises
Each exercise has a (rough) indication of its difficulty, as follows:
| * | easier: can be solved by (almost) only using relevant definitions/results, |
| ** | medium: in addition to relevant definitions/results, needs a limited amount of work/creativity, |
| *** | harder: in addition to relevant definitions/results, needs a larger amount of work/serious creativity, |
| 💀 | warning: might make your brain hurt! These are mainly intended to provide some extra challenge for those of you keen on that and are generally quite hard. You don't need to worry about these too much for exam purposes. |
The exam consists of mostly ** and *** level questions, some *, and possibly at most a few marks worth of 💀.
A bit of preaching: it is an incredibly important part of the study process to try and work on the exercises as much as possible. To become a better mathematician/learn new maths (and also to get a good exam mark ;)), above all you need to do it. And yes, of course that includes falling over things, and making mistakes, and getting stuck, and getting frustrated — all part of the game and what you’re supposed to be doing! Your lecturers have done that as well and still do it. What matters is that you don’t let that discourage you and that you make good use of the help and resources available to help you develop your skills. As part of that, many exercises have a hint in a block like this:
Hint!
These are trying to help you on your way if you don’t know where to start or to provide some ideas if you get stuck. In spirit of the above, always have a look at these first and try again before you look at the full solution. (These hints are an extra service that won’t be available in the exam I’m afraid ;).)
Of course, we have our classes and there’s office hours, email etc. as well — I’m at any time very happy to help you with any questions you may have, and you should please never feel that any question is “too dumb” to ask!
Full/detailed solutions for the exercises will become available, just immediately below the exercises, immediately after our tutorial hour (you may have to refresh the page).
Exercise 6.1 [*/**] Consider the risk \(X \sim \text{Unif}(0,100)\).
- Show that the premium for \(X\) based on the Variance premium principle with loading factor \(0.5\) amounts to \(1400/3\).
- Why could a premium of \(1400/3\) for \(X\) be considered unreasonable?
For part i, recall that the Variance premium principle was defined in Section 6.2.1, and that the mean and variance can simply be found in Appendix B!
For part ii, what is the range of \(X\) i.e. its set of possible values? If you’re not entirely sure, apply Equation A.18 (and if you’re not sure what the pdf of \(X\) is, recall that that guy is listed in Appendix B as well). How does the premium computed in part i compare to this range, and which desirable property listed in Section 6.2.2 is quite obviously violated?
For part i, from the attached distribution list (cf. Appendix B) we can read that \[\E[X]=\frac{0+100}{2}=50 \quad \text{and} \quad \var(X)=\frac{100^2}{12}=\frac{2500}{3}\] and hence, using the Variance premium principle from Section 6.2.1 we get that \[\pi(X)=\E[X]+0.5 \var(X)=\frac{1400}{3}.\]
For part ii, observe that the range of \(X\) i.e. its set of possible values is the interval \((0,100)\) — if you’re not sure: recall Equation A.18 and we can deduce from the attached distribution list (cf. Appendix B) that the pdf \(f_X\) of \(X\) satisfies \(f_X(x)>0\) if and only if \(x \in (0,100)\). Since \(1400/3>100\), this means that the premium is larger than the largest value that \(X\) can take! This is unreasonable because it violates the “No rip-off” property discussed in Section 6.2.2, see also the discussion there for the reasons why a reasonable premium should satisfy this property.
Exercise 6.2 [**] Let \(X\) be a risk that follows an \(\text{Exp}(\alpha)\) distribution, for some \(\alpha>0\). Compute the premium for \(X\) according to each of the five premium principles we discussed in Section 6.2.1.
Note: in the Exponential and Esscher principles you will see some infinities pop up for certain combinations of values for \(\alpha\) and the parameter \(\beta\) used in the def of these principles. In such cases, we consider the premium principle at hand simply to be undefined.
Using the goodies from Appendix B, the first three principles are just a matter of plugging stuff into their expressions. For the Exponential and Esscher principle, you’ll need to work out \(\E[e^{\beta X}]\) and \(\E[X e^{\beta X}]\). Use the standard approach of writing these in integral form (using Equation A.16, and pick up the pdf of \(X\) from Appendix B if needed). Be careful for the traps/inifnities here! If you’re unsure, consider the cases \(\alpha<\beta\), \(\alpha=\beta\) and \(\alpha>\beta\) separately — in the former two cases you will get infinities! For the integral for \(\E[X e^{\beta X}]\), a cute way to do it quick is to manipulate that integral towards an integral of the pdf of a suitable Gamma distribution since such an integral simplifies to \(1\) (cf. Equation A.13).
Let’s work through the list from Section 6.2.1 one by one. Recall (using Appendix B if needed) that we have avaiable that \(\E[X]=1/\alpha\) and \(\var(X)=1/\alpha^2\).
Expected Value principle: for any \(\beta \geq 0\) \[\pi(X)=(1+\beta) \E[X]=\frac{1+\beta}{\alpha}.\]
Variance principle: for any \(\beta \geq 0\) \[\pi(X)=\E[X]+\beta \var(X)=\frac{1}{\alpha}+\frac{\beta}{\alpha^2}.\]
Standard deviation principle: for any \(\beta \geq 0\) \[\pi(X)=\E[X]+\beta \sqrt{\var(X)}=\frac{1+\beta}{\alpha}.\]
Exponential principle: for any \(\beta > 0\) \[\pi(X)=\frac{1}{\beta} \log \big( \E[e^{\beta X}] \big). \tag{6.15}\] This is where we actually have to start doing some work, since our Appendix B unfortunately does not give us an expression for \(\E[e^{\beta X}]\)! Quick word of warning: since \(X\) is non-negative (indeed its range/set of possible values is \((0,\infty)\) — if you’re unsure then, as also pointed out in Exercise 6.1, recall Equation A.18 and note that the pdf \(f_X\) of \(X\) listed in Appendix B has the property that \(f_X(x)>0\) if and only if \(x>0\)), also the random variables \(e^{\beta X}\) and \(X e^{\beta X}\) are non-negative (including the latter for use in the Esscher principle below). That means that their expectation is well defined but it may equal \(\infty\)!
Now let’s work out \(\E[e^{\beta X}]\). Fix some \(\beta>0\). Using the standard integral formula (cf. Equation A.16) with the choice \(h(x)=e^{\beta x}\) as well as the pdf \(f_X\) of \(X\) (pick it up from Appendix B if needed) we get that \[\begin{align*} \E[e^{\beta X}] &= \E[h(X)] \\ &= \int_{-\infty}^\infty h(x) f_X(x) \d x \\ &= \int_0^\infty e^{\beta x} \alpha e^{-\alpha x} \d x \\ &= \alpha \int_0^\infty e^{(\beta-\alpha) x} \d x. \end{align*} \tag{6.16}\] Now be careful with this integral! If \(\beta=\alpha\) then we’re integrating the constant function \(1\) from \(x=0\) to \(x=\infty\) and hence the result is \(\infty\). If \(\beta \not= \alpha\) then we can work it out as \[\left. \frac{1}{\beta-\alpha} e^{(\beta-\alpha) x} \right|_0^\infty\] which equals \(1/(\alpha-\beta)\) if \(\beta-\alpha<0\) i.e. \(\alpha>\beta\) while it equals \(\infty\) if \(\beta-\alpha>0\) i.e. \(\alpha<\beta\). So altogether we get that \[\E[e^{\beta X}]=\begin{cases} \frac{\alpha}{\alpha-\beta} & \text{if } \alpha>\beta \\ \infty & \text{if } \alpha \leq \beta. \end{cases} \tag{6.17}\] Now plugging this back into Equation 6.15 we ultimately arrive at the following conclucion. If \(\alpha>\beta\) then \[\pi(X)=\frac{1}{\beta} \log \left( \frac{\alpha}{\alpha-\beta} \right),\] while if \(\alpha \leq \beta\) then \(\pi(X)\) is not well defined (or infinite, whichever you prefer).
Esscher principle: for any \(\beta > 0\) \[\pi(X)=\frac{\E[X e^{\beta X}]}{\E[e^{\beta X}]}. \tag{6.18}\] Bit more work for this one! We have worked out \(\E[e^{\beta X}]\) above already, and we conluded that it is only finite if \(\alpha>\beta\). In the other case it is infinite and is Equation 6.18 not sensibly defined. So we’ll stick with the case \(\alpha>\beta\). We still have something to do here: work out \(\E[X e^{\beta X}]\). We follow the standard route, as also used above for \(\E[e^{\beta X}]\) to write \[\begin{align*} \E[X e^{\beta X}] &= \int_{-\infty}^\infty x e^{\beta x} f_X(x) \d x \\ &= \int_0^\infty x e^{\beta x} \alpha e^{-\alpha x} \d x \\ &= \alpha \int_0^\infty x e^{(\beta-\alpha) x} \d x \\ &= \alpha \int_0^\infty x e^{-(\alpha-\beta) x} \d x \\ &= \frac{\alpha}{(\alpha-\beta)^2} \int_0^\infty (\alpha-\beta)^2 x e^{-(\alpha-\beta) x} \d x \\ &= \frac{\alpha}{(\alpha-\beta)^2}, \end{align*} \] where in the final three equalities we cunningly work our integral towards the integral of the pdf of a \(\text{Gamma}(2,\alpha-\beta)\) distribution (cf. Appendix B, as also mentioned there, towards the bottom: \(\Gamma(2)=1!=1\)) since we know that that integral equals \(1\), cf. Equation A.13). Alternatively you could e.g. attack this integral using integration by parts.
So plugging this into Equation 6.18, also using Equation 6.17, we get that for \(\alpha>\beta\) \[\pi(X)=\frac{\alpha/(\alpha-\beta)^2}{\alpha/(\alpha-\beta)}=\frac{1}{\alpha-\beta}.\]
Note: if you find dealing with integrals as for the Exponential and Esscher principles above challenging, then consider the following strategy. By very definition, it holds that \[\int_0^\infty ... \d x = \lim_{b \to \infty} \int_0^b ... \d x.\] Using this you first fix an arbitrary \(b>0\) and work out \(\int_0^b ... \d x\). You won’t encounter any infinities there. Then take the limit for \(b \to \infty\), where then your infinities pop up in maybe a clearer way. For instance, for the integral we used to compute \(\E[e^{\beta X}]\) i.e. (cf. Equation 6.16) \[\int_0^\infty e^{(\beta-\alpha) x} \d x\] (ignoring the factor \(\alpha\) in front of it at the moment), you would first fix an arbitrary \(b>0\) and work out \[\int_0^b e^{(\beta-\alpha) x} \d x.\] You still have to be careful with the antiderivative, you’re naturally going to \[\int_0^b e^{(\beta-\alpha) x} \d x = \left. \frac{1}{\beta-\alpha} e^{(\beta-\alpha) x} \right|_0^b = \frac{1}{\beta-\alpha} \left( e^{(\beta-\alpha) b}-1 \right),\] but this is only valid if \(\beta \not= \alpha\) (indeed there is a very big hint in the observation that for \(\beta = \alpha\) you’re dividing by \(0\) in your antiderivative!). For \(\beta = \alpha\) we rather get that \[\int_0^b e^{(\beta-\alpha) x} \d x = \int_0^b 1 \d x=b.\] Then take the limit for \(b \to \infty\), and you see the infinity appearing in the second case as well as in the first case when \(\beta-\alpha>0\) i.e. \(\alpha<\beta\).
Exercise 6.3 [**/***] In Section 6.2.2 we discussed a number of desirable properties for premium principles, and in Example 6.2 we discussed some examples of when a premium principle does or does not satisfies such properties. Here we pursue some more of these.
- Show that the Standard Deviation premium principle does satisfy the Scale Invariance property, but not the Additivity property.
- Show that the Esscher premium principle satisfies both the Consistency and Additivity properties.
For part i, the Standard Deviation should be pretty straightforward (however don’t forget, what is \(\var(cX)\) again, something with a square somewhere eh) and to disprove the Additivity property, following the lead of Example 6.2, one option is to go and find explicit examples of random variables \(X\) and \(Y\) where the property fails. Consider for instance independent \(X,Y\) with \[X=\begin{cases} 0 & \text{with prob } 1/2 \\ a & \text{with prob } 1/2 \end{cases} \quad \text{and} \quad Y=\begin{cases} 0 & \text{with prob } 1/2 \\ b & \text{with prob } 1/2 \end{cases} \] for some \(a,b>0\).
For part ii, nothing really special is required, just work your way through the algebra from one side to the other, working our brackets, taking constants out the of the expectation etc. For the Additivity one, don’t forget about Equation A.56.
For part i, recall from Section 6.2.1 that the Standard Deviation principle is given by \[\pi(X)=\E[X]+ \beta \sqrt{\var(X)},\] where \(\beta \geq 0\) is the loading factor.
- The Scale Invariance property (cf. Section 6.2.2) states that for any \(X \in \mathcal{D}\) and constant \(c>0\) it holds that \(\pi(cX)=c \pi(X)\). Indeed this is satisfied: \[\begin{align*} \pi(cX) &= \E[cX]+ \beta \sqrt{\var(cX)} \\ &= c\E[X]+\beta \sqrt{c^2 \var(X)} \\ &= c \left( \E[X]+\beta \sqrt{\var(X)} \right) \\ &= c \pi(X), \end{align*} \] where we used the standard linearity rules for the expectation and variance.
- The Additivity property (cf. Section 6.2.2) states that for independent \(X,Y \in \mathcal{D}\), it holds that \(\pi(X+Y)=\pi(X)+\pi(Y)\). We’re asked to show that this property is not satisfied. There’s several ways to do this, but to stay in line with the general strategy suggested in Example 6.2 let’s follow that one again: try the simplest example(s) of random variable(s) that you can think of, and see if you can find that indeed the property does not hold for the choices you made. In this case we need two random variables, so let’s simply take for some \(a>0\) and \(b>0\) independent random variables \(X,Y\) given by \[X=\begin{cases} 0 & \text{with prob } 1/2 \\ a & \text{with prob } 1/2 \end{cases} \quad \text{and} \quad Y=\begin{cases} 0 & \text{with prob } 1/2 \\ b & \text{with prob } 1/2. \end{cases} \] Then we have that \(\E[X]=a/2\), \(\var(X)=a^2/4\), \(\E[Y]=b/2\) and \(\var(X)=b^2/4\) (cf. Example 6.2). Further \(\E[X+Y]=\E[X]+\E[Y]=(a+b)/2\) and, using the independence, \(\var(X+Y)=\var(X)+\var(Y)=(a^2+b^2)/4\). So \[\begin{align*} \pi(X+Y) &= \E[X+Y]+ \beta \sqrt{\var(X+Y)} \\ &=\frac{a+b}{2}+\beta \sqrt{\frac{a^2+b^2}{4}} \end{align*} \] and \[\begin{align*} \pi(X)+\pi(Y) &= \frac{a}{2}+\beta \sqrt{\frac{a^2}{4}}+\frac{b}{2}+\beta \sqrt{\frac{b^2}{4}} \\ &=\frac{a+b}{2}+\beta \left( \frac{a}{2}+\frac{b}{2} \right). \end{align*} \] So Additivity does not hold provided that \[\sqrt{\frac{a^2+b^2}{4}} \not= \frac{a}{2}+\frac{b}{2}. \tag{6.19}\] You could come up with a general argument, i.e. prove Equation 6.19 for any \(a,b>0\), but also don’t forget that all we need is one single example where Additivity breaks. So giving specific choices for \(a,b>0\) leading to Equation 6.19 is also perfectly fine. Take for instance \(a=b=1\), then the left hand side of Equation 6.19 equals \(\sqrt{2}/2\) while the right hand side equals \(1\), clearly unequal and so done!
Unexaminable note: if you are wondering how to construct the random variables \(X\) and \(Y\) that we used above, to ensure that they actually exist — we can always do this: consider two separate experiments/probability spaces, one on which \(X\) lives and one on which \(Y\) lives. Like rolling a die A as one experiment and a die B as another experiment. Then join both together into one experiment. In the case of the dice: roll both die A and die B. Mathematically speaking this is a product (probability) space.
For part ii, recall from Section 6.2.1 that the Esscher premium principle is given by \[\pi(X)=\frac{\E[X e^{\beta X}]}{\E[e^{\beta X}]},\] where \(\beta > 0\) is a parameter.
- The Consistency property (cf. Section 6.2.2) states that for any \(X \in \mathcal{D}\) and constant \(c>0\) it holds that \(\pi(X+c)=\pi(X)+c\). This is straightforward to check with some algebra: \[\begin{align*} \pi(X+c) &= \frac{\E[(X+c) e^{\beta (X+c)}]}{\E[e^{\beta (X+c)}]} \\ &= \frac{\E[e^{\beta c} X e^{\beta X}+e^{\beta c} c e^{\beta X}]}{\E[e^{\beta c} e^{\beta X}]} \\ &= \frac{e^{\beta c} \E[X e^{\beta X}+ c e^{\beta X}]}{e^{\beta c} \E[ e^{\beta X}]} \\ &= \frac{\E[X e^{\beta X}+ c e^{\beta X}]}{\E[ e^{\beta X}]} \\ &= \frac{\E[X e^{\beta X}]+ c \E[e^{\beta X}]}{\E[ e^{\beta X}]} \\ &= \frac{\E[X e^{\beta X}]}{\E[e^{\beta X}]}+c \\ &= \pi(X)+c. \end{align*} \]
- The Additivity property (cf. Section 6.2.2) states that for independent \(X,Y \in \mathcal{D}\), it holds that \(\pi(X+Y)=\pi(X)+\pi(Y)\). Again a bit of algebra really: \[\begin{align*} \pi(X+Y) &= \frac{\E[(X+Y) e^{\beta (X+Y)}]}{\E[e^{\beta (X+Y)}]} \\ &= \frac{\E[X e^{\beta X} e^{\beta Y}+Y e^{\beta X} e^{\beta Y}]}{\E[e^{\beta X} e^{\beta Y}]} \\ &= \frac{\E[X e^{\beta X}] \E[e^{\beta Y}]+\E[Y e^{\beta X}] \E[e^{\beta Y}]}{\E[e^{\beta X}] \E[e^{\beta Y}]} \\ &= \frac{\E[X e^{\beta X}]}{\E[e^{\beta X}]} + \frac{\E[Y e^{\beta Y}]}{\E[e^{\beta Y}]} \\ &= \pi(X)+\pi(Y), \end{align*} \] where in the third equality we used that an expectation of products becomes a product of expectations due to the independence of \(X\) and \(Y\) (this is an application of Equation A.56).
Exercise 6.4 [**] Time for some comps with risk measures!
- Suppose that \(X\) is a discrete random variable with range/possible values \(\{1,2,3\}\), and write the distribution of \(X\) as \[\P(X=1)=p, \quad \P(X=2)=q, \quad \P(X=3)=1-p-q\] for some \(p,q \geq 0\) satisfying \(p+q \leq 1\). For what value(s) of \(p\) and \(q\) do we have that \(\VaR_{0.9}(X)=2\)?
- Suppose that \(X \sim \text{Exp}(\theta)\) for some \(\theta>0\). For some level \(p \in (0,1)\), work out the Value-at-Risk \(\VaR_p(X)\) and the Conditional Tail Expectation \(\cte_p(X)\).
- Now suppose that \(X\) has a Pareto distribution with pdf Equation 6.12. For some level \(p \in (0,1)\), work out the Value-at-Risk \(\VaR_p(X)\) and the Conditional Tail Expectation \(\cte_p(X)\) (warning: infinities are lurking!).
For part i: first convince yourself, using Definition 6.4 (in particular the paragraph predecing it) and/or Example 6.3, that \(\VaR_{0.9}(X)=2\) if and only if \(\P(X>2) \leq 0.1\) and \(\P(X>1) > 0.1\). Then translate this to conditions on \(p\) and \(q\).
For part ii: the VaR is in principle a straightforward application of Definition 6.4, you’ll need the inverse cdf of course but we’ve seen (almost) that guy previously, look for instance at Example 2.3 if you’d like some help with it. For the CTE it’s again a matter of invoking its def (cf. Definition 6.5) and working it out. If you need some inspiration, have a look at Example 6.4 for instance. You should get to \[\cte_p(X) = \frac{1}{1-p} \int_{-\log(1-p)/\theta}^\infty x\theta e^{-\theta x} \d x,\] and this integral can be slaughtered using integration by parts for instance.
For part iii: use the same ideas and steps as in part ii. For the CTE you should get \[\cte_p(X) = \frac{1}{1-p} \int_{(1-p)^{-1/\alpha}}^\infty \alpha x^{-\alpha} \d x.\] Look out for the special case \(\alpha=1\) as well as for values of \(\alpha\) where this integral is infinite.
For part i: Recall from Definition 6.4 (in particular the paragraph predecing it) and/or Example 6.3 that \(\VaR_{0.9}(X)\) equals the smallest \(k \in \{1,2,3\}\) for which \(\P(X>k) \leq 1-0.9=0.1\). So \(\VaR_{0.9}(X)=2\) if and only if \(\P(X>2) \leq 0.1\) and \(\P(X>1) > 0.1\). Since \[\P(X>2)=\P(X=3)=1-p-q \quad \text{and} \quad \P(X>1)=1-\P(X=1)=1-p,\] we have \(\VaR_{0.9}(X)=2\) if and only if \(1-p-q \leq 0.1\) and \(1-p>0.1\), which we can rewrite to \(p<0.9\) and \(q \geq 0.9-p\).
For part ii: For the VaR, following the same arguments as in Example 2.3 (but now \(\theta\) is not set to \(1\) as it was there!) we can first recall — or work out from the pdf (recall from Appendix B) and Equation A.12 that the cdf \(F_X\) is given by \[F_X(x)=\begin{cases} 1-e^{-\theta x} & \text{if } x>0 \\ 0 & \text{if } x \leq 0. \end{cases} \] If you sketch this guy, you quickly realise that for every \(y \in (0,1)\) fixed, there is a unique solution \(x\) to the equation \(F_X(x)=y\) which can be found as follows: \[F_X(x)=y \iff 1-e^{-\theta x}=y \iff x=\frac{-1}{\theta} \log(1-y).\] Hence the (classic in this case) inverse cdf \(F_X^{-1}: (0,1) \to \R\) is given by \[F_X^{-1}(y)=\frac{-1}{\theta} \log(1-y) \quad \text{for all } y \in (0,1).\] Now we can apply Definition 6.4 to write down that \[\VaR_p(X)=F_X^{-1}(p)=\frac{-1}{\theta} \log(1-p). \tag{6.20}\]
Next for the CTE. Starting from Definition 6.5 the first few steps are standard and essentially same as in Example 6.4 for the \(\text{Unif}(0,1)\) distribution: \[\begin{align*} \cte_p(X) &= \E \big[X \, \big| \, X>\VaR_p(X) \big] \\ &= \frac{\E \left[\mathbf{1}_{\{X>-\log(1-p)/\theta\}} X \right]}{\P(X>\VaR_p(X))} \\ &= \frac{1}{1-p} \int_{-\log(1-p)/\theta}^\infty x\theta e^{-\theta x} \d x, \end{align*} \tag{6.21}\] where for the second equality we used Equation A.43 and plugged in Equation 6.20, and for the third Equation A.45 with \(I=(-\log(1-p)/\theta,\infty)\) as well as that by very def \(\P(X>\VaR_p(X))=1-p\) (recall from Definition 6.4).
The integral in Equation 6.21 is a bit annoying, but nothing a bit of integration by parts can’t take care of (as also used in Example 6.1 for instance): \[\begin{align*} \int_{-\log(1-p)/\theta}^\infty x\theta e^{-\theta x} \d x &= -x e^{-\theta x} \big|_{x=-\log(1-p)/\theta}^{x=\infty} + \int_{-\log(1-p)/\theta}^\infty e^{-\theta x} \d x \\ &= -\frac{\log(1-p)}{\theta} e^{\log(1-p)} + \left. \frac{-1}{\theta} e^{-\theta x} \right|_{x=-\log(1-p)/\theta}^{x=\infty} \\ &= -\frac{\log(1-p)}{\theta} (1-p) + \frac{1}{\theta} (1-p) \\ &= \frac{1-p}{\theta} \big( 1-\log(1-p) \big). \end{align*} \] Plugging this back into Equation 6.21, we arrive at \[\cte_p(X)=\frac{\big( 1-\log(1-p) \big)}{\theta}.\]
For part iii: For the VaR, we start again by finding the cdf. Using Equation A.12 and the pdf from Equation 6.12 we can easily compute \[\begin{align*} F_X(x) &= \int_{-\infty}^x f_X(y) \d y \\ &= \begin{cases} 0 & \text{if } x \leq 1 \\ \int_1^x \alpha y^{-\alpha-1} \d y = 1-x^{-\alpha} & \text{if } x > 1. \end{cases} \end{align*} \] Again a quick sketch shows that for every \(y \in (0,1)\) fixed, there is a unique solution \(x\) to the equation \(F_X(x)=y\), found as follows: \[F_X(x)=y \iff 1-x^{-\alpha}=y \iff x=(1-y)^{-1/\alpha}.\] Hence the (again classic) inverse cdf \(F_X^{-1}: (0,1) \to \R\) is given by \[F_X^{-1}(y)=(1-y)^{-1/\alpha} \quad \text{for all } y \in (0,1)\] and from Definition 6.4: \[\VaR_p(X)=F_X^{-1}(p)=(1-p)^{-1/\alpha}.\] For the CTE, using the same arguments as used in Equation 6.21 we know get that \[\begin{align*} \cte_p(X) &= \E \big[X \, \big| \, X>\VaR_p(X) \big] \\ &= \frac{\E \left[\mathbf{1}_{\{X>(1-p)^{-1/\alpha}\}} X \right]}{\P(X>\VaR_p(X))} \\ &= \frac{1}{1-p} \int_{(1-p)^{-1/\alpha}}^\infty x \alpha x^{-\alpha-1} \d x \\ &= \frac{1}{1-p} \int_{(1-p)^{-1/\alpha}}^\infty \alpha x^{-\alpha} \d x. \\ \end{align*} \tag{6.22}\] For this integral, note that for \(\alpha=1\) we can use \(\log(x)\) as antiderivative and the result is \(\infty\). For any \(\alpha>0\), \(\alpha \not= 1\) \[\begin{align*} \int_{(1-p)^{-1/\alpha}}^\infty \alpha x^{-\alpha} \d x &= \left. \frac{\alpha}{1-\alpha} x^{1-\alpha} \right|_{x=(1-p)^{-1/\alpha}}^{x=\infty} \\ &= \begin{cases} \infty & \text{if } \alpha<1 \\ \frac{\alpha}{\alpha-1} (1-p)^{(\alpha-1)/\alpha} & \text{if } \alpha>1. \end{cases} \end{align*} \] If you get tangled up in the infinities, then don’t forget the alternative route we also mentioned in the sol of Exercise 6.2, namely to write \[\int_{(1-p)^{-1/\alpha}}^\infty \alpha x^{-\alpha} \d x = \lim_{b \to \infty} \int_{(1-p)^{-1/\alpha}}^b \alpha x^{-\alpha} \d x.\]
Finally then, plugging this all back into Equation 6.22 we then finally arrive at \[ \cte_p(X) = \begin{cases} \infty & \text{if } \alpha \leq 1 \\ \frac{\alpha}{\alpha-1} (1-p)^{(\alpha-1)/\alpha-1}=\frac{\alpha}{\alpha-1} (1-p)^{-1/\alpha} & \text{if } \alpha>1. \end{cases} \]
Exercise 6.5 [💀] This question is not very hard and actually very interesting and insightful I think (especially for Actuarial Science students and any others among you potentially interested in this direction!), however I shouldn’t overload you so it is not “core” and you shouldn’t worry about it for exam purposes. We illustrate here how the positive loading property from Section 6.2.2 very naturally arises, especially in the context of a collection of many risks of identical nature as is the typical habitat of an insurance company for instance.
Simple forms of insurance are around for thousands of years already, in the form of a collective of individuals who bring together many relatively small amounts of money (the premiums) in a fund which is then used to support the few individuals who suffer financial losses that would otherwise completely bankrupt them. A simple form of social insurance if you will.
Among the earliest known forms is the following. Imagine sea merchants, many centuries ago, who made a living by sailing the seas in their ships to purchase goods in some part of the world and sell them on in some other part. Think: ships made of wood facing potentially very rough conditions at sea, not much in the way of reliable weather predictions yet, pirates, mutiny, you name it — a whole list of things that could go horribly wrong.
His ship was a merchant’s financially most valuable possession: requiring big investment and if it got lost at sea then it typically meant a life of debts and poverty ahead, for themselves and/or for their families if they themselves didn’t make it back.
Against this background an insurance fund was established: every merchant taking part in it would pay in a premium and in return would get financial assistance if something happened to their ship. In the terminology of Section 6.1: ownership of the risk of damage to the ship is transferred from the merchant to the insurance fund, in exchange for that premium payment.
Let’s make this more precise (while keeping it simple): say that each merchant owns one ship, that the worth of ship plus any good it carries is £1,000, and that each ship gets lost with probability \(0.01\). So the risk \(X\) resulting from the potential loss of a ship is a random variable with the following distribution: \[X=\begin{cases} 1{,}000 & \text{with prob } 0.01 \\ 0 & \text{with prob } 0.99. \end{cases} \tag{6.23}\]
Suppose that there are \(n\) merchants taking part in the fund, and let \(X_1, \ldots, X_n\) denote their individual risks (i.e. each of these random variables is a copy of \(X\)). We assume that these random variables are independent. Note that the total risk/loss (or: aggregate risk as we will call it in the upcoming Chapter 7 & Chapter 8) is the random variable \(S\) given by \[S=X_1+\ldots+X_n. \tag{6.24}\] We say that the fund gets ruined if its total loss is larger than the total amount of premiums it has collected — clearly in this case it won’t be able to live up to its purpose of taking care of all the merchants losing their ships.
Explain why we may equally well write \(S=1{,}000 N\), where \(N \sim \text{Binomial}(n,0.01)\).
Suppose that each merchant pays a premium of £15. Find the probability that the fund gets ruined in three different cases: when there are \(n=65\) merchants, when there are \(n=1{,}000\) merchants and finally when there are \(n=10{,}000\) merchants.
Hint: recall the
pbinom()function in R from Section 1.4.4.Repeat part ii, but now under the assumption that each merchant pays a premium of £10.
Repeat part ii one more time, but now under the assumption that each merchant pays a premium of £5.
- Can you explain how the results of parts ii–iv motivate that the fund should make sure to respect the positive loading property from Section 6.2.2?
- Finally consider a more general setting, where we allow the risk/loss for each merchant to be a random variable \(X\) with any (non-negative) distribution as long as it has finite mean and variance. Using the Central Limit Theorem (cf. Section A.7), investigate the importance of the fund charging a premium satisfying the positive loading property from Section 6.2.2 for large \(n\).
- Let’s wrap up with a blast from the past (namely our copulas discussions from Chapter 3–Chapter 5): do you think that the assumption we have used throughout, namely that the individual risks of the merchants are independent, is reasonable?
For part i, what is the interpreation of a Binomial distribution again? Something with repeating the same experiment a certain number of times and counting the number of “successes” was it?
For parts ii–iv, compute the total income of the fund (number of merchants times premium per merchant) and use that it gets ruined if \(S\) is larger than this total income. Using the expression from part i you can translate this into a probability for \(N\) which you can then easily compute, e.g. using the pbinom() function in R as the hint points out.
For part v, note that parts ii–iv seem to indicate that there are three different regimes: one for the premium larger than £10, one for the premium equal to £10, and one for the premium less than £10. Which regime(s) is best for the fund? And how does that £10 relate to \(\E[X]\)?
For part vi, say that \(n\) (large) merchants pay a premium of \(P\) each, yielding a total income of \(nP\). The total outgo/loss is \(S=X_1+\ldots+X_n\), where the \(X_i\)’s are independent copies of \(X\). Note that the probability of ruin for the fund equals \(\P(S>nP)\). Use the Central Limit Theorem (cf. Section A.7) to express this prob (approximately) in terms of the cdf of a standard Normal distribution, and then use its properties (in particular Equation A.5) to analyse this prob for the three cases \(P<\E[X]\), \(P=\E[X]\) and \(P>\E[X]\).
For part vii, open question! :).
For part i, two possible arguments:
- More mathematical: we can express Equation 6.23 as \(X=1{,}000 I\), where \[I=\begin{cases} 1 & \text{with prob } 0.01 \\ 0 & \text{with prob } 0.99 \end{cases} \] i.e. \(I \sim \text{Bernoulli}(0.01)\) (cf. Appendix B if needed). This means that we can write Equation 6.24 also as \[S=1{,}000 I_1+\ldots+1{,}000 I_n=1{,}000 (I_1+\ldots+I_n),\] and you’ll remember from an earlier prob course that the sum of \(n\) independent \(\text{Bernoulli}(p)\) distributed random variables has a \(\text{Binomial}(n,p)\) distribution.
- More intutive: let \(N\) denote the number of ships out of the \(n\) in total that gets lost. Then we’re looking at executing \(n\) independent and identical experiments, each of which results in “success” (=ship gets lost) with prob \(0.01\). Then as we know, the total number of “successes” \(N\) follows a \(\text{Binomial}(n,0.01)\) distribution. And since each ship lost costs the fund £1,000, their total risk/loss is \(S=1{,}000 N\).
For part ii, as follows:
- Let \(n=65\). Then the total premium income of the fund is \(65 \cdot 15=975\). Since each lost ship costs the fund \(1{,}000\), the fund gets ruined if at least one ship gets lost i.e. with probability \[\P(N \geq 1)=1-\P(N=0)=1-0.99^{65} \approx 0.48.\] That’s a pretty large probability for such a disastrous event!
- Now let \(n=1{,}000\). Then the total premium income of the fund is \(1{,}000 \cdot 15=15{,}000\). So the fund gets ruined if 16 or more ships get lost, which happens with probability (cf. Listing 6.2 below) \[\P(N \geq 16)=1-\P(N \leq 15) \approx 1-0.95 \approx 0.05.\] That’s considerably smaller than the \(0.48\) we found for \(n=65\)!
- Finally let \(n=10{,}000\). Then the total premium income of the fund is \(10{,}000 \cdot 15=150{,}000\). So the fund gets ruined if 151 or more ships get lost, which happens with probability (cf. Listing 6.2 below) \[\P(N \geq 151)=1-\P(N \leq 150) \approx 1-0.9999989 \approx 0.000001,\] That’s extremely unlikely!
So the trend seems pretty clear: if the merchants pay a premium of £15, then the more merchants sign up the safer the fund is.
For part iii, the only thing that changes in comparison with part ii is the total premium income. If you adjust the comps in part ii for that, you should find the following probabilities of ruin for the fund:
- \(n=65\): \(\approx 0.48\) (unchanged)
- \(n=1{,}000\): \(\approx 0.42\)
- \(n=10{,}000\): \(\approx 0.48\)
So now we see a completely different picture: the probability of ruin seems to hoovering around \(\approx 0.5\) no matter what value of \(n\).
For part iv, again a change in premium only. You should now find the following probabilities of ruin for the fund:
- \(n=65\): \(\approx 0.48\) (unchanged)
- \(n=1{,}000\): \(\approx 0.93\)
- \(n=10{,}000\): \(\approx 1.00\)
So yet again a completely different picture: the probability of ruin now seems to explode very quickly as \(n\) grows, towards even almost certain ruin for \(n=10{,}000\)!
For part v, if we “guesstrapolate” a bit from parts ii–iv, then a sensible guess would be that the premium £10 forms something of a critial amount: if the premium is larger than £10 then the fund is quite safe as long as enough merchants take part, and the more the better; while if the premium is less than £10 then the fund will never be very safe and in fact the more merchants take part, the more likely it is that disaster strikes, even essentially certain if \(n\) is large enough.
So how can we explain that magical amount of £10? Well indeed, it is exactly the expectation of each individual merchant risk, cf. Equation 6.23: \[\E[X]=1{,}000 \cdot 0.01 + 0 \cdot 0.99=10.\] So if we are right with our “guesstrapolation” here (and we indeed are, as you could confirm by doing more experimenting for other premium amounts and/or use part vi below), then it is absolutely crucial for the fund to charge each merchant a premium that is at least equal to \(\E[X]\) (and ideally a bit more). And that is of course exactly what the positive loading property from Section 6.2.2 entails! :).
Intuitively we could for instance think about this as follows. If a large \(n\) number of merchants takes part, then by the Law of Large Numbers, the expected outgo for the fund is roughly \(n\) times the expected loss per merchant. If they all pay a premium that is (somewhat) larger than their individual expected loss, then the income for the fund can be written as their expected outgo plus \(n\) times (premium minus individual expected loss). That second term forms a safety margin on top of their expected outgo which grows with \(n\). Hence the larger \(n\), the more unlikely it is that the actual outgo spikes so far above the expected outgo that it also eats up that safety margin and actually gets the fund into trouble. This is made more precise in part vi below.
For part vi, suppose that each of the (large) \(n\) merchant pays a premium of \(P\). Then the total income of the fund is \(nP\). At the outgo side, the total loss \(S\) given by \[S=X_1+\ldots+X_n,\] where the \(X_i\)’s are i.i.d. copies of the risk per merchant \(X\). Let’s write \(\mu_X:=\E[X]\) and \(\sigma_X^2:=\var(X)\). Then by standard arguments (recall e.g. from Equation A.57) \[\E[S]=n \mu_X \quad \text{and} \quad \var(S)=n \sigma_X^2.\]
Now, the fund gets ruined if and only if the total loss is larger than the total income i.e. if \(S>nP\), and the probability that this happens is \[\begin{align*} \P(S>nP) &= 1-\P(S \leq nP) \\ &= 1-\P \left( \frac{S-\E[S]}{\sqrt{\var(S)}} \leq \frac{nP-\E[S]}{\sqrt{\var(S)}} \right) \\ &= 1-\P \left( \frac{S-\E[S]}{\sqrt{\var(S)}} \leq \frac{nP-n\mu_X}{\sigma_X \sqrt{n}} \right) \\ &\approx 1-\Phi \left( \frac{nP-n\mu_X}{\sigma_X \sqrt{n}} \right) \\ &= 1-\Phi \left( \frac{\sqrt{n}}{\sigma_X} (P-\mu_X) \right), \end{align*} \tag{6.25}\] where the fourth equality uses the CLT (cf. Section A.7) and \(\Phi\) is the cdf of a standard Normal distribution.
Indeed, recalling that \(\Phi\) is a cdf and hence is a non-decreasing function with limit values (cf. Equation A.5) \[\lim_{x \to -\infty} \Phi(x)=0 \quad \text{and} \quad \lim_{x \to -\infty} \Phi(x)=1,\] we can deduce from Equation 6.25 exactly the behaviour we also already saw in parts ii–iv:
- if \(P<\mu_X\) (as in part iv), then \(\sqrt{n} (P-\mu_X)/\sigma_X\) is a large negative number, and hence the probability that ruin happens i.e. \(\P(S>nP)\) is quite close to \(1\);
- if \(P=\mu_X\) (as in part iii), then \(\sqrt{n} (P-\mu_X)/\sigma_X=0\) and hence the probability that ruin happens is \(\P(S>nP)=1-\Phi(0)=1/2\);
- if \(P>\mu_X\) (as in part ii), then \(\sqrt{n} (P-\mu_X)/\sigma_X\) is a large positive number, and hence the probability that ruin happens i.e. \(\P(S>nP)\) is quite close to \(0\).
So, clearly the fund wants to be in a situation where at least \(P \geq \mu_X\) and ideally \(P >\mu_X\) holds, i.e. where the positive loading property holds!
For part vii, it’s hard to draw firm conclusions without more specific info. An argument could be that as long as the merchants are mostly active in different parts of the world, then the independence assumption seems reasonable. However if they all tend to follow the same shipping route and/or have the same harbour as their departure point for instance, then there is the risk of a catastrophic event (such as a large storm along the shared route or a large fire in the harbour for example) causing losses for many merchants simultanuously.
Don’t underestimate how much of a difference this can make (as we also, in general terms, extensively discussed in Chapter 7 & Chapter 8)! Take for instance the setting of part ii above, with a premium of £15 per merchant and \(n=10{,}000\) merchants taking part. As we saw there, under the independence asumption the fund is virtually certain not to get ruined. However the collected premiums are only sufficient to cover 150 lost ships. So if there is some probability \(p\) that some catastrophic event happens in which more than 150 ships are lost then the probability that the fund gets ruined drastically changes to at least \(p\) as well!