Conditional Probabilities
摘要
Starting from a discrete probability space \((\Omega,{\mathcal{P}}(\Omega),\mathbb{P})\) and a set \(B\subseteq\Omega\) with \(\mathbb{P}(B)> 0\) , one obtains another probability measure on \({\mathcal{P}}(\Omega)\) via \(\displaystyle\mathbb{P}^{B}:{\mathcal{P}}(\Omega)\to[0,1],\quad A\mapsto\frac{\mathbb{P}(A\cap B)}{\mathbb{P}(B)}.\) Since \(\mathbb{P}^{B}(B)=1\) , one interprets the probability \(\mathbb{P}^{B}(A)\) as the probability of the event A given that the event B definitely occurs (conditional probability). One has, so to speak, reduced the set of possible outcomes Ω to the set B. If now \(\mathbb{P}^{B}(A)=\mathbb{P}(A)\) applies, it follows that \(\mathbb{P}(A\cap B)=\mathbb{P}(A)\cdot\mathbb{P}(B)\) ; in this case, the probability for A is not affected by the reduction of the set of outcomes from Ω to B; one says, the events A and B are stochastically independent. Therefore, when introducing the function I to measure the amount of information, we also demanded \(\displaystyle I(pq)=I(p)+I(q);\) if the “message” A occurs with probability p, the message B with probability q and both messages with probability pq, then the respective amounts of information should add up if both messages occur (since stochastic independence is present).