<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>statistics &#8211; Twenty Third Floor</title>
	<atom:link href="https://twentythirdfloor.co.za/category/statistics/feed/" rel="self" type="application/rss+xml" />
	<link>https://twentythirdfloor.co.za</link>
	<description>Perspectives</description>
	<lastBuildDate>Wed, 17 Dec 2025 09:33:44 +0000</lastBuildDate>
	<language>en-GB</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=6.9.1</generator>

<image>
	<url>https://twentythirdfloor.co.za/blog_files/wp-content/uploads/2011/07/cropped-cropped-IMG_5265_2-2-32x32.jpg</url>
	<title>statistics &#8211; Twenty Third Floor</title>
	<link>https://twentythirdfloor.co.za</link>
	<width>32</width>
	<height>32</height>
</image> 
	<item>
		<title>Why did your premium go up when someone else hit your parked car?</title>
		<link>https://twentythirdfloor.co.za/2025/12/17/why-did-your-premium-go-up-when-someone-else-hit-your-parked-car/</link>
					<comments>https://twentythirdfloor.co.za/2025/12/17/why-did-your-premium-go-up-when-someone-else-hit-your-parked-car/#respond</comments>
		
		<dc:creator><![CDATA[David Kirk]]></dc:creator>
		<pubDate>Wed, 17 Dec 2025 09:33:44 +0000</pubDate>
				<category><![CDATA[Actuarial and Risk]]></category>
		<category><![CDATA[complexity]]></category>
		<category><![CDATA[insurance]]></category>
		<category><![CDATA[measurement]]></category>
		<category><![CDATA[statistics]]></category>
		<guid isPermaLink="false">https://twentythirdfloor.co.za/?p=3204</guid>

					<description><![CDATA[It feels unfair. You did nothing wrong. Someone else drove into your car outside your house and now you&#8217;re paying more. Surely that&#8217;s just the insurer clawing back their loss? Maybe. But maybe not. Let me take the scenic route to explaining why. Driving my daughter to school, I occasionally spot someone tailgating or cutting [&#8230;]]]></description>
										<content:encoded><![CDATA[
<p>It feels unfair. You did nothing wrong. Someone else drove into your car outside your house and now you&#8217;re paying more. Surely that&#8217;s just the insurer clawing back their loss?</p>



<p>Maybe. But maybe not. Let me take the scenic route to explaining why.</p>



<p>Driving my daughter to school, I occasionally spot someone tailgating or cutting across a solid line. What&#8217;s striking is how often within seconds I&#8217;ll see them make a second and third thoughtless move. Weaving without indicating, pushing into the off-ramp queue at the last moment &#8211; all the sorts of things nobody does in the queue at Woolies when there is accountability. Ring of Gyges stuff &#8211; and frankly unfair that both types of people populate this universe.</p>



<p>Yes, there&#8217;s confirmation bias here. I&#8217;m sure I make mistakes that annoy others. Attribution bias too &#8211; we forgive our own lapses as innocent mistakes while judging others harshly. I don&#8217;t think that&#8217;s all of it though.</p>



<p>What we&#8217;re observing is that observations are usually not independent. Errors cluster. They share root causes, be it time pressure, risk tolerance, cell phone use, attitudes towards others, personality, upbringing and more. Whatever produces one lapse doesn&#8217;t reset between intersections. One observation carries information about an underlying factor you can&#8217;t directly see.</p>



<p>When you see a pattern, update your priors on what comes next.</p>



<p>Back to that not-at-fault claim. A claim happened. And empirically, one claim predicts the next. The insurer often doesn&#8217;t know why, but the signal is there in the data. Maybe your car is parked on a busy road. Maybe it&#8217;s closer to a corner than ideal, or near a distracting intersection. Maybe street lighting is poor. Maybe it&#8217;s just a high-traffic area where the probability of someone having a bad day near your vehicle is elevated. As more complex analysis tools (yes including AI&#8230;) are applied, with more data and more context, we may get closer to understanding the hidden risk factors underneath the high level outcomes. Either way, the correlation is there.</p>



<p>The customer likely did nothing wrong. The premium increase can still reflect rational Bayesian updating.</p>



<p>Now, is it always this principled? No. Sometimes it&#8217;s crude loss-ratio management. The insurer just wants to recover what they paid out over time. Sometimes the pricing model can&#8217;t even distinguish at-fault from not-at-fault, so everything gets the same treatment. These practices exist, and they&#8217;re harder to defend.</p>



<p>The legitimate version, where claims predict future claims for reasons the policyholder can&#8217;t fully observe or control, that&#8217;s real too. And if the signal is real, <em>not</em> adjusting the premium means other policyholders cross-subsidise the risk. The question isn&#8217;t whether someone pays, but who.</p>



<p>If you feel this is still unfair, well I think you&#8217;re correct on some level, albeit not yet a practical one:</p>



<p>All risk rating involves a choice about what &#8220;fair&#8221; means. Is it fair that you pay a price tailored to your risk, even if that risk stems from factors you didn&#8217;t choose? Or is it fairer that we all pay the same, pooling our luck and misfortune together? The young driver pays more not because they decided to be nineteen, but because nineteen-year-olds crash more often. Risk-based fairness says differentiate. Solidarity-based fairness says pool.</p>



<p>These aren&#8217;t the same definition, and better maths won&#8217;t reconcile them. We try to draw lines, some factors &#8220;feel&#8221; acceptable, others don&#8217;t, but those lines are social and political choices, not mathematical ones. Therefore they will differ between people, between firms, and over time.</p>



<p>So next time your premium moves in a way that feels unjust, ask whether the insurer is being lazy or seeing something real. But also ask who else would pay if you didn&#8217;t. Insurance is a group exercise, and the maths doesn&#8217;t care about fault. Only we do.</p>
]]></content:encoded>
					
					<wfw:commentRss>https://twentythirdfloor.co.za/2025/12/17/why-did-your-premium-go-up-when-someone-else-hit-your-parked-car/feed/</wfw:commentRss>
			<slash:comments>0</slash:comments>
		
		
			</item>
		<item>
		<title>The Perils of Value-at-Risk and Portfolio Insurance</title>
		<link>https://twentythirdfloor.co.za/2024/12/09/the-perils-of-value-at-risk-and-portfolio-insurance/</link>
					<comments>https://twentythirdfloor.co.za/2024/12/09/the-perils-of-value-at-risk-and-portfolio-insurance/#respond</comments>
		
		<dc:creator><![CDATA[David Kirk]]></dc:creator>
		<pubDate>Mon, 09 Dec 2024 07:00:00 +0000</pubDate>
				<category><![CDATA[Actuarial and Risk]]></category>
		<category><![CDATA[capital]]></category>
		<category><![CDATA[financial risk]]></category>
		<category><![CDATA[managing uncertainty]]></category>
		<category><![CDATA[market risk]]></category>
		<category><![CDATA[measurement]]></category>
		<category><![CDATA[statistics]]></category>
		<guid isPermaLink="false">https://twentythirdfloor.co.za/?p=3075</guid>

					<description><![CDATA[It is essential to consider critical viewpoints that challenge conventional wisdom—especially when it comes to Value-at-Risk (VaR). In a thought-provoking dialogue, Nassim Taleb critiques VaR and highlights the dangers of portfolio insurance and dynamic hedging strategies. Here are key arguments from his 1997 forceful response to Philippe Jorion’s support for VaR. Misplaced Precision and Concrete [&#8230;]]]></description>
										<content:encoded><![CDATA[
<p><mark style="background-color:rgba(0, 0, 0, 0)" class="has-inline-color has-primary-color"><strong>It is essential to consider critical viewpoints that challenge conventional wisdom—especially when it comes to Value-at-Risk (VaR).</strong></mark></p>



<p>In a thought-provoking dialogue, Nassim Taleb critiques VaR and highlights the dangers of portfolio insurance and dynamic hedging strategies. Here are key arguments from his 1997 forceful response to Philippe Jorion’s support for VaR.</p>



<h4 class="wp-block-heading">Misplaced Precision and Concrete Metrics</h4>



<p>Taleb warns that the unique precision of VaR creates a false sense of certainty. He describes this as a form of &#8220;misplaced concreteness,&#8221; where risk managers mistakenly believe they have a comprehensive understanding of potential losses based solely on point estimates. This can lead to dangerous oversimplifications in risk assessment, potentially masking the underlying complexities of market behavior.</p>



<h4 class="wp-block-heading">The Risks of Portfolio Insurance</h4>



<p>Dynamic hedging, often employed in portfolio insurance, is particularly perilous. Taleb argues that these strategies can exacerbate market downturns, relying on flawed statistical models that underestimate tail risks. When events occur that fall outside expected parameters, the repercussions can be catastrophic, as seen in past financial crises where such strategies failed to provide the intended safety net.</p>



<h4 class="wp-block-heading">Standard Error vs. Point Estimates</h4>



<p>A critical issue Taleb raises is the phenomenon where the standard error of a risk estimate can exceed the estimate itself. This stark mismatch reveals the inherent dangers of relying on these calculations. The history of financial crises shows that bizarrely improbable events—deemed unlikely by VaR—frequently materialize, often with devastating consequences that could have been better anticipated with a more qualitative understanding of risk.</p>



<h4 class="wp-block-heading">Forecasting Volatility</h4>



<p>Taleb emphasizes that accurately forecasting volatility is exceptionally challenging. The reliance on historical data and models leads to a blind spot regarding unpredictable market dynamics. This difficulty only compounds the risks associated with tools like VaR and portfolio insurance, which may provide a false sense of security in the face of uncertainty.</p>



<h4 class="wp-block-heading">The Illusion of Credibility</h4>



<p>Moreover, the widespread adoption of VaR among financial institutions is not a measure of scientific credibility. Instead, it often reflects a collective oversight of significant risks, leading to disastrous outcomes. Financial institutions may become overly reliant on VaR, neglecting other qualitative assessments of risk that could better inform their strategies.</p>



<p>While we are all familiar with George Box&#8217;s quote, &#8220;All models are wrong, but some are useful,&#8221; Taleb&#8217;s perspective might be paraphrased more pessimistically: &#8220;All models are wrong, and most are downright dangerous.&#8221; This insight serves as a crucial reminder that while models can aid in decision-making, they are not infallible and should not be the sole basis for risk management.</p>



<h4 class="wp-block-heading">Conclusion</h4>



<p>For 2025, let&#8217;s all recognise the limitations of our tools and the potential pitfalls of over-reliance on quantitative metrics. By fostering a deeper understanding of risk through both quantitative and qualitative lenses, we can better prepare for the unpredictable nature of financial markets.</p>



<p>ðŸ”— Explore the full discussion for deeper insights: <a href="https://www.fooledbyrandomness.com/jorion.html">Nassim Taleb Replies to Philippe Jorion, 1997</a></p>
]]></content:encoded>
					
					<wfw:commentRss>https://twentythirdfloor.co.za/2024/12/09/the-perils-of-value-at-risk-and-portfolio-insurance/feed/</wfw:commentRss>
			<slash:comments>0</slash:comments>
		
		
			</item>
		<item>
		<title>Capital Modelling for parametric insurance &#8211; intro</title>
		<link>https://twentythirdfloor.co.za/2024/10/21/capital-modelling-for-parametric-insurance-intro/</link>
					<comments>https://twentythirdfloor.co.za/2024/10/21/capital-modelling-for-parametric-insurance-intro/#respond</comments>
		
		<dc:creator><![CDATA[David Kirk]]></dc:creator>
		<pubDate>Mon, 21 Oct 2024 09:01:56 +0000</pubDate>
				<category><![CDATA[Actuarial and Risk]]></category>
		<category><![CDATA[alternative investments]]></category>
		<category><![CDATA[capital]]></category>
		<category><![CDATA[capital structure]]></category>
		<category><![CDATA[climate change]]></category>
		<category><![CDATA[insurance]]></category>
		<category><![CDATA[managing uncertainty]]></category>
		<category><![CDATA[measurement]]></category>
		<category><![CDATA[microinsurance]]></category>
		<category><![CDATA[modelling]]></category>
		<category><![CDATA[Solvency Assessment and Management]]></category>
		<category><![CDATA[Solvency II]]></category>
		<category><![CDATA[statistics]]></category>
		<guid isPermaLink="false">https://twentythirdfloor.co.za/?p=3063</guid>

					<description><![CDATA[As parametric insurance gains traction, insurers face specific challenges in capital modeling and regulatory capital navigation. I have a longer paper coming out on this, but if you&#8217;re looking for an intro, here are some of the interesting and different aspects compared to more traditional insurance. 1. Regulatory Uncertainty: The treatment of parametric insurance under [&#8230;]]]></description>
										<content:encoded><![CDATA[
<p>As parametric insurance gains traction, insurers face specific challenges in capital modeling and regulatory capital navigation. I have a longer paper coming out on this, but if you&#8217;re looking for an intro, here are some of the interesting and different aspects compared to more traditional insurance.<br /><br />1. <strong>Regulatory Uncertainty</strong>: The treatment of parametric insurance under frameworks like Solvency II and SAM remains ambiguous. Insurers must engage proactively with regulators to establish appropriate methodologies. Regulators have the challenge of how to shoe-horn parametric insurance into a regulatory framework that was not designed with this in mind. For example, in South Africa, a of 2024 at least, parametric non-life insurance is approved on  case by case basis under a regulatory sandbox, but as &#8220;non insurance business&#8221;.  This is because under current regulations, &#8220;non life insurance&#8221; must be on an indemnity basis.<br /><br />2. <strong>Line of Business Allocation</strong>: Fitting parametric products into traditional lines of business is complex. Many parametric products resemble inwards non-proportional reinsurance more than direct insurance, with payouts triggered by specific events. Even then, there is no guarantee that the standard premium volatility factors are appropriate. Insurers may need to explore Undertaking/Insurer Specific Parameters (USP / ISP) or transition to partial internal models. For now, this &#8220;non insurance business&#8221; approved in South Africa has typically been allocated to the agriculture LoB for capital purposes. This may match the nature of the business (typically drought or rainfall related) but there is no reason to believe that the variability in claims will match that of other agricultural business. I wonder whether &#8220;inwards non proportional reinsurance&#8221; might be a better fit in some ways. The reserve risk parameters will hopefully be too conservative &#8211; since the a key idea behind parametric insurance is very quick and objective claim settlement without extended reporting or payment delays.<br /><br />3. <strong>Portfolio Size and Trigger Remoteness</strong>: The risk profile changes significantly with smaller portfolio sizes and trigger remoteness. As triggers become more remote, the capital required relative to premium increases. At a certain point, the 99.5th VaR can fall well outside the 3-sigma range, challenging standard deviation-based approaches. <br /><br />4. <strong>Diversification Effects</strong>: Understanding correlation between parametric triggers, and at different levels of triggers, means approaches like copula modeling might be necessary. Student t copulas are a likely candidate.  As portfolios grow and become more diversified this may moderate. However, there will almost always be fewer sensors / indices than individual policyholders and risk exposures. Therefore I expect challenges on diversification to continue.<br /><br />5. <strong>Attritional vs. Catastrophic Losses</strong>: The binary nature of parametric triggers blurs the line between attritional and catastrophic losses. <br /><br />6. <strong>Time Series vs. One-Year Capital View</strong>: While sensor data forms a time series that could be modeled using techniques like SARIMAX or GARCH-X, the one-year capital view required by regulations doesn&#8217;t necessarily need to incorporate this time series structure. The complex physics-based models that are increasingly used for pricing and prediction will likely remain too unwieldy for capital purposes for an extended period.<br /><br />7. <strong>Climate risk and trends</strong>: An advantage of parametric insurance is the typical clean time-series sensor records (necessary for pricing and risk management). However, the continued relevance of historical records is at risk given climate change for many key parametric coverages.<br /><br />8. <strong>Demonstrating Appropriateness</strong>: The Head of Actuarial Function (HAF) faces the challenge of demonstrating that the chosen capital approach appropriately reflects the risk profile of parametric products. The approach needs to work within the regulatory framework, but the result must still be reasonable. </p>



<figure class="wp-block-image size-large"><a href="https://twentythirdfloor.co.za/blog_files/wp-content/uploads/2024/10/image.png"><img fetchpriority="high" decoding="async" width="1024" height="273" src="https://twentythirdfloor.co.za/blog_files/wp-content/uploads/2024/10/image-1024x273.png" alt="" class="wp-image-3065" srcset="https://twentythirdfloor.co.za/blog_files/wp-content/uploads/2024/10/image-1024x273.png 1024w, https://twentythirdfloor.co.za/blog_files/wp-content/uploads/2024/10/image-300x80.png 300w, https://twentythirdfloor.co.za/blog_files/wp-content/uploads/2024/10/image-768x204.png 768w, https://twentythirdfloor.co.za/blog_files/wp-content/uploads/2024/10/image.png 1093w" sizes="(max-width: 1024px) 100vw, 1024px" /></a></figure>



<p><br /><br />As the parametric insurance market evolves, so too must our approach to capital modeling. The challenges are significant, but so are the opportunities for innovation and more accurate risk assessment.</p>
]]></content:encoded>
					
					<wfw:commentRss>https://twentythirdfloor.co.za/2024/10/21/capital-modelling-for-parametric-insurance-intro/feed/</wfw:commentRss>
			<slash:comments>0</slash:comments>
		
		
			</item>
		<item>
		<title>Why gender at birth is not a coin toss</title>
		<link>https://twentythirdfloor.co.za/2019/08/26/why-gender-at-birth-is-not-a-coin-toss/</link>
					<comments>https://twentythirdfloor.co.za/2019/08/26/why-gender-at-birth-is-not-a-coin-toss/#respond</comments>
		
		<dc:creator><![CDATA[David Kirk]]></dc:creator>
		<pubDate>Mon, 26 Aug 2019 12:25:45 +0000</pubDate>
				<category><![CDATA[Actuarial and Risk]]></category>
		<category><![CDATA[complexity]]></category>
		<category><![CDATA[future studies]]></category>
		<category><![CDATA[news]]></category>
		<category><![CDATA[statistics]]></category>
		<guid isPermaLink="false">https://twentythirdfloor.co.za/?p=2744</guid>

					<description><![CDATA[I love analogies. As analogies go, many things are well approximated by the &#8220;toss of a coin&#8221;. Gender at birth is not. And it&#8217;s not for two reasons, only one of which you can probably guess. The sex ratio at birth is not 1 boy to 1 girl. As of 2017, according to World Bank [&#8230;]]]></description>
										<content:encoded><![CDATA[
<p>I love analogies. As analogies go, many things are well approximated by the &#8220;toss of a coin&#8221;. Gender at birth is not. And it&#8217;s not for two reasons, only one of which you can probably guess.</p>



<p>The sex ratio at birth is not 1 boy to 1 girl.  As of 2017, <a href="https://data.worldbank.org/indicator/SP.POP.BRTH.MF?most_recent_value_desc=false&amp;view=chart">according to World Bank data, it was 1.073 male births for every 1 female birth</a>.  In my mind, I still operate on the heuristic of 1.05:1.   (It was 1.062:1 in 1962, which probably dates some of my professors rather than me.  The US is currently at 1.048:1.)</p>



<p>There are several postulated reasons, including generally higher mortality of males, but also under-registration of female births, and sex-selective abortion and infanticide of female babies. Both of those last two could be attributed to China&#8217;s one-child policy.  China&#8217;s ratio is still 1.15:1 compared to Sierra Leone at 1.018.  (Zimbabwe is close to bottom at 1.02 with South Africa not much greater at 1.03.)</p>



<p>The second reason gender at birth is not a coin toss is less well known. The gender of a child is not independent of the gender of the previous child. The impact is small and best-studied on (non-human) animals. One thought is that the &#8220;condition&#8221; of the mother has an impact on the gender of the child. Other possible factors include frequency of intercourse and paternal age.</p>



<span id="more-2744"></span>



<p>Since many of these factors are not independent, the probability of same-gender children is slightly higher than a pure coin toss would indicate, even if it were a biased 1:073:1 coin.</p>



<p>This post was prompted by the interesting analysis of 12 consecutive girl births in a tiny village in Poland. The &#8220;statistics lecturer&#8221; who explained quite well why this isn&#8217;t actually a surprise glossed over some, interesting to me, and arguably important distinctions that demographers (and parents) care about perhaps a little more.</p>



<p>The next Really Interesting Question is what happens to sex ratios, and births and genetic characteristics and selection for children in the coming decades. I can easily imagine that my own children will be the last in &#8220;my line&#8221; to be all-natural, unedited, unmodified, naturally selected genetic offspring.</p>
]]></content:encoded>
					
					<wfw:commentRss>https://twentythirdfloor.co.za/2019/08/26/why-gender-at-birth-is-not-a-coin-toss/feed/</wfw:commentRss>
			<slash:comments>0</slash:comments>
		
		
			</item>
		<item>
		<title>More statistical fun from XKCD</title>
		<link>https://twentythirdfloor.co.za/2013/07/11/more-statistical-fun-from-xkcd/</link>
					<comments>https://twentythirdfloor.co.za/2013/07/11/more-statistical-fun-from-xkcd/#respond</comments>
		
		<dc:creator><![CDATA[David Kirk]]></dc:creator>
		<pubDate>Thu, 11 Jul 2013 08:34:15 +0000</pubDate>
				<category><![CDATA[Actuarial and Risk]]></category>
		<category><![CDATA[statistics]]></category>
		<guid isPermaLink="false">http://twentythirdfloor.co.za/?p=2235</guid>

					<description><![CDATA[XKCD is always good, but the statistical references are often quite insightful!]]></description>
										<content:encoded><![CDATA[<p>XKCD is always good, but the statistical references are often quite insightful!</p>
<p><a href="http://xkcd.com/1236/"><img decoding="async" class="alignnone" alt="" src=" http://imgs.xkcd.com/comics/seashell.png" width="286" height="393" /></a></p>
]]></content:encoded>
					
					<wfw:commentRss>https://twentythirdfloor.co.za/2013/07/11/more-statistical-fun-from-xkcd/feed/</wfw:commentRss>
			<slash:comments>0</slash:comments>
		
		
			</item>
		<item>
		<title>The virtual irrelevancy of population size to required sample size</title>
		<link>https://twentythirdfloor.co.za/2013/04/05/the-virtual-irrelevancy-of-population-size-to-required-sample-size/</link>
					<comments>https://twentythirdfloor.co.za/2013/04/05/the-virtual-irrelevancy-of-population-size-to-required-sample-size/#respond</comments>
		
		<dc:creator><![CDATA[David Kirk]]></dc:creator>
		<pubDate>Thu, 04 Apr 2013 22:38:01 +0000</pubDate>
				<category><![CDATA[Actuarial and Risk]]></category>
		<category><![CDATA[data analysis]]></category>
		<category><![CDATA[insight]]></category>
		<category><![CDATA[managing uncertainty]]></category>
		<category><![CDATA[measurement]]></category>
		<category><![CDATA[statistics]]></category>
		<guid isPermaLink="false">http://twentythirdfloor.co.za/?p=2170</guid>

					<description><![CDATA[Statistics and sampling are fundamental to almost all of our understanding of the world. The world is too big to measure directly. Measuring representative samples is a way to understand the entire picture. Popular and academic literature are both full of examples of poor sample selection resulting in flawed conclusions about the population. Some of [&#8230;]]]></description>
										<content:encoded><![CDATA[<p>Statistics and sampling are fundamental to almost all of our understanding of the world. The world is too big to measure directly. Measuring representative samples is a way to understand the entire picture.</p>
<p>Popular and academic literature are both full of examples of poor sample selection resulting in flawed conclusions about the population. Some of the most famous examples relied on sampling from telephone books (in the days when phone books still mattered and only relatively wealthy people had telephones) resulting in skewed samples.</p>
<p>This post is not about bias in sample selection but rather the simpler matter of sample sizes.</p>
<h2>Population size is usually irrelevant to sample size</h2>
<p>I&#8217;ve read too often the quote: &#8220;Your sample was only 60 people from a population of 100,000.Â  That&#8217;s not statistically relevant.&#8221;Â  Which is of course plain wrong and frustratingly wide-spread.</p>
<p>Required Sample Size is dictated by:</p>
<ul>
<li>How accurate one needs the estimate to be</li>
<li>The standard deviation of the population</li>
<li>The homogeneity of the population</li>
</ul>
<p><strong>Only in exceptional circumstances does population size matter at all</strong>. To demonstrate this, consider the graph of the standard error of the mean estimate as the sample size increases for a population of 1,000 with a standard deviation of the members of the population of 25.</p>
<p><figure id="attachment_2172" aria-describedby="caption-attachment-2172" style="width: 563px" class="wp-caption alignnone"><a href="https://twentythirdfloor.co.za/blog_files/wp-content/uploads/2013/04/Sample-Sized-P1000.png"><img decoding="async" class="size-full wp-image-2172" alt="Standard Error as Sample Size increases for population of 1,000" src="https://twentythirdfloor.co.za/blog_files/wp-content/uploads/2013/04/Sample-Sized-P1000.png" width="563" height="392" srcset="https://twentythirdfloor.co.za/blog_files/wp-content/uploads/2013/04/Sample-Sized-P1000.png 563w, https://twentythirdfloor.co.za/blog_files/wp-content/uploads/2013/04/Sample-Sized-P1000-300x208.png 300w, https://twentythirdfloor.co.za/blog_files/wp-content/uploads/2013/04/Sample-Sized-P1000-430x300.png 430w" sizes="(max-width: 563px) 100vw, 563px" /></a><figcaption id="caption-attachment-2172" class="wp-caption-text">Standard Error as Sample Size increases for population of 1,000</figcaption></figure></p>
<p>The standard error drops very quickly at first, then decreases very gradually thereafter even for a large sample of 100. Let&#8217;s see how this compares to a larger population of 10,000.<span id="more-2170"></span></p>
<p><figure id="attachment_2173" aria-describedby="caption-attachment-2173" style="width: 563px" class="wp-caption alignnone"><a href="https://twentythirdfloor.co.za/blog_files/wp-content/uploads/2013/04/Sample-Sized-P10000.png"><img loading="lazy" decoding="async" class="size-full wp-image-2173" alt="Standard Error as Sample Size increases for population of 1,000 vs 10,000" src="https://twentythirdfloor.co.za/blog_files/wp-content/uploads/2013/04/Sample-Sized-P10000.png" width="563" height="392" srcset="https://twentythirdfloor.co.za/blog_files/wp-content/uploads/2013/04/Sample-Sized-P10000.png 563w, https://twentythirdfloor.co.za/blog_files/wp-content/uploads/2013/04/Sample-Sized-P10000-300x208.png 300w, https://twentythirdfloor.co.za/blog_files/wp-content/uploads/2013/04/Sample-Sized-P10000-430x300.png 430w" sizes="auto, (max-width: 563px) 100vw, 563px" /></a><figcaption id="caption-attachment-2173" class="wp-caption-text">Standard Error as Sample Size increases for population of 1,000 vs 10,000</figcaption></figure></p>
<p>It&#8217;s not an error that those graphs are almost identical.Â  (see next section for why they&#8217;re not exactly the same.) <strong>Population size isÂ  irrelevant.</strong></p>
<h2>Finite Population Correction &#8211; or when population size begins to matter</h2>
<p>Population size is usually irrelevant to sample size. There is an exception. Few people even know about this because it is so rarely relevant.</p>
<p>If one samples the entire population, there is no room for error in estimating statistics related to the population. The mean of the population is then known precisely. We don&#8217;t have a &#8220;sample&#8221; any more. The standard error of our estimate is thus zero since it is no longer an estimate but a direct measurement of the population statistics.</p>
<p><figure id="attachment_2174" aria-describedby="caption-attachment-2174" style="width: 563px" class="wp-caption alignnone"><a href="https://twentythirdfloor.co.za/blog_files/wp-content/uploads/2013/04/Sample-Sized-P100.png"><img loading="lazy" decoding="async" class="size-full wp-image-2174" alt="Standard Error as Sample Size increases for population of 100 vs 1,000" src="https://twentythirdfloor.co.za/blog_files/wp-content/uploads/2013/04/Sample-Sized-P100.png" width="563" height="392" srcset="https://twentythirdfloor.co.za/blog_files/wp-content/uploads/2013/04/Sample-Sized-P100.png 563w, https://twentythirdfloor.co.za/blog_files/wp-content/uploads/2013/04/Sample-Sized-P100-300x208.png 300w, https://twentythirdfloor.co.za/blog_files/wp-content/uploads/2013/04/Sample-Sized-P100-430x300.png 430w" sizes="auto, (max-width: 563px) 100vw, 563px" /></a><figcaption id="caption-attachment-2174" class="wp-caption-text">Standard Error as Sample Size increases for population of 100 vs 1,000</figcaption></figure></p>
<p>As the sample size approaches the population size, the sample error begins to decline. We adjust the normal standard error estimate by multiplying it by sqrt((Populaton &#8211; Sample) / (Population &#8211; 1)). In this case, you can see that the 100 population sample error does begin to decrease in the graph above.</p>
<p>You can go ahead and forget that formula now as you&#8217;ll probably never need it. It does explain the slight difference in the graphs in the previous section though.</p>
<h2>Non-homogenous populations aka sample size still isn&#8217;t the answer</h2>
<p>One last point &#8211; if the population is not homoegenous, this will typically increase the total population variability and result in higher standard errors. It adds risk in using a small sample size. The best solution is not to simply increase the population size. Cluster Sampling is a straightforward way of reducing the variability of overall estimates for population statistics and better understanding differences between the clusters.</p>
<h2>So what determines the required sample size?</h2>
<p>I get asked &#8220;how big a sample will we need&#8221; typically with only the population size to go on. Hopefully I&#8217;ve explained why the population size is irrelevant in almost all scenarios. How should you estimate the required sample size?</p>
<p>Ideally, you need an initial estimate of population standard deviation. You can get this from prior studies of similar populations or by doing a quick, small initial survey to estimate the standard error to set a sample size. You could also continue to sample under the standard error of your estimate became sufficiently low, but this often isn&#8217;t practical and can encourage unintended biases in terms of how the sample is increased over time.</p>
<p>Without anything to go on, I use the rule of thumb that 10 observations is usually the minimum number within a homogenous population to measure a single characteristic. 20 is better and usually, more doesn&#8217;t make much of a difference.</p>
<p>Smaller than you thought?</p>
]]></content:encoded>
					
					<wfw:commentRss>https://twentythirdfloor.co.za/2013/04/05/the-virtual-irrelevancy-of-population-size-to-required-sample-size/feed/</wfw:commentRss>
			<slash:comments>0</slash:comments>
		
		
			</item>
		<item>
		<title>XKCD and statisticians</title>
		<link>https://twentythirdfloor.co.za/2013/03/29/xkcd-and-statisticians/</link>
					<comments>https://twentythirdfloor.co.za/2013/03/29/xkcd-and-statisticians/#comments</comments>
		
		<dc:creator><![CDATA[David Kirk]]></dc:creator>
		<pubDate>Fri, 29 Mar 2013 05:00:45 +0000</pubDate>
				<category><![CDATA[statistics]]></category>
		<guid isPermaLink="false">http://twentythirdfloor.co.za/?p=2162</guid>

					<description><![CDATA[I was looking for an XKCD comic to insert into a presentation on insurance regulations. Yes, I do that. I came across a comic I hadn&#8217;t seen before and that also illustrates a fundamentally important point. Enjoy.]]></description>
										<content:encoded><![CDATA[<p>I was looking for an <a href="http://www.xkcd.com/">XKCD </a>comic to insert into a presentation on insurance regulations. Yes, I do that. I came across a comic I hadn&#8217;t seen before and that also illustrates a fundamentally important point.</p>
<p>Enjoy.</p>
<p><img loading="lazy" decoding="async" class="alignnone" alt="" src="http://imgs.xkcd.com/comics/frequentists_vs_bayesians.png" width="468" height="709" /></p>
]]></content:encoded>
					
					<wfw:commentRss>https://twentythirdfloor.co.za/2013/03/29/xkcd-and-statisticians/feed/</wfw:commentRss>
			<slash:comments>1</slash:comments>
		
		
			</item>
		<item>
		<title>Birdy statistics</title>
		<link>https://twentythirdfloor.co.za/2013/03/22/birdy-statistics/</link>
					<comments>https://twentythirdfloor.co.za/2013/03/22/birdy-statistics/#respond</comments>
		
		<dc:creator><![CDATA[David Kirk]]></dc:creator>
		<pubDate>Fri, 22 Mar 2013 07:30:22 +0000</pubDate>
				<category><![CDATA[measurement]]></category>
		<category><![CDATA[statistics]]></category>
		<guid isPermaLink="false">http://twentythirdfloor.co.za/?p=2148</guid>

					<description><![CDATA[I&#8217;m not sure about this fascinating article on birds evolving to avoid cars in the US. The story is that fewer cliff swallows are being killed on the roads AND those birds killed have longer than average wings. The argument here is that longer wings make for less agility, making the birds more likely to [&#8230;]]]></description>
										<content:encoded><![CDATA[<p>I&#8217;m not sure about this <a href="http://green.24.com/swallows-evolve-to-dodge-cars/">fascinating article on birds evolving to avoid cars in the US</a>.</p>
<p>The story is that fewer cliff swallows are being killed on the roads AND those birds killed have longer than average wings. The argument here is that longer wings make for less agility, making the birds more likely to be killed by cars.</p>
<p>So far so good. But then:</p>
<blockquote><p>The authors of the study found that over a 30 year period, annual cliff swallow roadkill has declined steadily from 20 birds per season in 1984 and 1985 to less than five birds per season during the last five years. Over the same period, traffic volumes remained the constant and the overall bird populations increased.</p></blockquote>
<p>I am not an ornithologist or evolutionary expert, but I just can&#8217;t see how between 20 and 5 birds killed per season will create enough selection pressure to change the wingspan.</p>
<p>The <a href="http://download.cell.com/current-biology/pdf/PIIS0960982213001942.pdf?intermediate=true">original research</a> summary is far more persuasive than the article. It shows graphs and statistical test results for decreasing average population wing size and increasing average road-kill wing size over time.</p>
<p>The explanation of why the average wingspan for cliff swallows killed be vehicles should increase is left unexplained. It does rather suggest potential measurement or confirmation bias from the research team &#8211; once the hypothesis starts looking interesting it would be very easy to unintentionally bias the measurements.Â  Measuring wingspan accurately to within a few millimetres is fraught with risks of subjective error.</p>
<p>Further, it looks like around 3 data points contribute significantly to the low p values of the tests and I would be very curious to know how robust the results were to removal of these influential points. It looks like the trends might remain, but without anything close to the significant suggested by the original research.</p>
<p>Finally, the clustering of wing measurement points in certain years suggests different levels of care and accuracy in measurement and potential &#8220;anchoring and adjustment bias&#8221;. It&#8217;s very hard to apply the same measurement protocols over 30 years.</p>
<p>So, fascinating research, interesting conclusion, but I&#8217;m left somehow unconvinced. It&#8217;s a pity the statistics applied weren&#8217;t a little more robust and the obvious criticisms weren&#8217;t addressed.</p>
<p>&nbsp;</p>
]]></content:encoded>
					
					<wfw:commentRss>https://twentythirdfloor.co.za/2013/03/22/birdy-statistics/feed/</wfw:commentRss>
			<slash:comments>0</slash:comments>
		
		
			</item>
		<item>
		<title>US Life Expectancy and the dangers of superficial analysis</title>
		<link>https://twentythirdfloor.co.za/2012/11/16/us-life-expectancy-and-the-dangers-of-superficial-analysis/</link>
					<comments>https://twentythirdfloor.co.za/2012/11/16/us-life-expectancy-and-the-dangers-of-superficial-analysis/#respond</comments>
		
		<dc:creator><![CDATA[David Kirk]]></dc:creator>
		<pubDate>Fri, 16 Nov 2012 12:39:10 +0000</pubDate>
				<category><![CDATA[Actuarial and Risk]]></category>
		<category><![CDATA[healthcare]]></category>
		<category><![CDATA[life insurance]]></category>
		<category><![CDATA[measurement]]></category>
		<category><![CDATA[statistics]]></category>
		<guid isPermaLink="false">http://twentythirdfloor.co.za/?p=2039</guid>

					<description><![CDATA[Life Expectancy is going up. In general. But what really matters isn&#8217;t the general but the specifics. I know it&#8217;s hard to work through maths and actual calculations, but it doesn&#8217;t help if you run your analysis off slogans. US Life Expectancy is going up. But not as much &#8220;at retirement&#8221; as it is up [&#8230;]]]></description>
										<content:encoded><![CDATA[<p>Life Expectancy is going up. In general. But what really matters isn&#8217;t the general but the specifics. I know it&#8217;s hard to work through maths and actual calculations, but it doesn&#8217;t help if you run your analysis off slogans.</p>
<p><a href="http://theincidentaleconomist.com/wordpress/zombie-life-expectancy-arguments/">US Life Expectancy is going up. But not as much &#8220;at retirement&#8221; as it is up &#8220;at birth&#8221; because all the improvements in infant mortality are irrelevant at retirement. Similarly, mortality improvements aren&#8217;t the same for all income bands.</a> The detail matters when it comes to trillions of dollars of social security.</p>
]]></content:encoded>
					
					<wfw:commentRss>https://twentythirdfloor.co.za/2012/11/16/us-life-expectancy-and-the-dangers-of-superficial-analysis/feed/</wfw:commentRss>
			<slash:comments>0</slash:comments>
		
		
			</item>
		<item>
		<title>More on gender and analytical problems</title>
		<link>https://twentythirdfloor.co.za/2012/07/31/more-on-gender-and-analytical-problems/</link>
					<comments>https://twentythirdfloor.co.za/2012/07/31/more-on-gender-and-analytical-problems/#respond</comments>
		
		<dc:creator><![CDATA[David Kirk]]></dc:creator>
		<pubDate>Tue, 31 Jul 2012 21:54:27 +0000</pubDate>
				<category><![CDATA[complexity]]></category>
		<category><![CDATA[insight]]></category>
		<category><![CDATA[investments]]></category>
		<category><![CDATA[measurement]]></category>
		<category><![CDATA[statistics]]></category>
		<guid isPermaLink="false">http://twentythirdfloor.co.za/?p=1895</guid>

					<description><![CDATA[Here is another interesting story with a gender angle. A study shows that stockholders in companies with women in the Board achieved better returns than those without. The obvious and likely correct point is that women add something valuable to the Board and is the company performs better. Diversity is a good thing in general, [&#8230;]]]></description>
										<content:encoded><![CDATA[<p>Here is another interesting story with a gender angle. A <a href="http://www.slate.com/blogs/moneybox/2012/07/31/companies_with_women_on_their_boards_perform_better.html">study shows that stockholders in companies with women in the Board achieved better returns than those without</a>.<br />
The obvious and likely correct point is that women add something valuable to the Board and is the company performs better. Diversity is a good thing in general, not least when it comes to considering complex issues with multiple stakeholders. It makes a good deal of sense to get this result. </p>
<p>Of course it&#8217;s not the only possible reason. It&#8217;s also not absolute proof that putting women onto a men-only Board would improve performance. </p>
<p>The problem is cause and effect. It might be that enlightened Boards add well-performing companies are more likely to add women to their Board. It might be that successful companies spend the time to get their Board composition right.</p>
<p>Finally, it might be stronger than the diversity argument. Women may simply be better at running companies than men. It&#8217;s a pity there are too few women-only boards to compare their performance to help answer the question whether women are better board members than men or is the benefit simply one of diversity. Interesting implications for other forms of diversity on Boards too.</p>
]]></content:encoded>
					
					<wfw:commentRss>https://twentythirdfloor.co.za/2012/07/31/more-on-gender-and-analytical-problems/feed/</wfw:commentRss>
			<slash:comments>0</slash:comments>
		
		
			</item>
	</channel>
</rss>
