The chaotic collection of ever-growing quantities of disparate data about businesses, society and individuals regularly attracts comment in the press. Insurance has not been spared. Indeed, the rise of Big Data has fostered the myth of an insurer capable of exploiting this mass of high-value data to determine each policyholder's individual behaviour in real time. Segmentation taken to extremes would allow premiums to be tailored precisely to each person, satisfying consumers who demand to pay "the" price corresponding exactly to their level of risk. On 2 October 2014, the magazine Les Échos asserted, without hesitation or embarrassment: "How Big Data will revolutionize insurance." The newspaper Le Monde was more cautious, asking in a headline on 22 November 2016: "Can Big Data really set insurance prices?" The academic world was less categorical still: writing in La Tribune on 20 January 2015, Jean-Pascal Gayant preferred the conditional mood when considering the future of solidarity. In short, risk pooling would have to be reinvented to account for the emergence of predictive models with almost infallible powers of anticipation.
Pricing: obtaining data
But what is actually the case? Mathematics can help us answer that question. First, what exactly do pricing procedures involve? An insurance contract guarantees a beneficiary the payment of specified sums if particular events—claims—occur, provided that the policyholder pays a premium set in advance (ex ante) by the insurer. This premium is determined from a body of objective information collected before the contract is signed.