4  Gummy bears

Last modified on 08. October 2026 at 09:03:09

“Gummy Bears! Bouncing here and there and everywhere; High adventure that’s beyond compare; They are the Gummy Bears.” — Gummy Bears Theme, Adventures of the Gummy Bears

This chapter provides a brief interlude to present an analysis of a dataset that I am personally connected to. If you work with data that you have generated, that data becomes very special. It’s like the rose in The Little Prince. It is not just one of thousands of data sets. It is my data set. Like every other data set, this one has a special story behind it. Where there is a story, there are also emotions. Follow me on the path to truth, guided by gummy bears.

4.1 General background

Like any good research story, this one begins in a gay club in Berlin with mirrored bathrooms, where you can see everything. Nothing stays hidden. When you can observe everything, no questions remain unanswered. This is somehow boring and not worthwhile for a truth-seeker. However, if you cannot see the truth before you, then you must draw conclusions based on your incomplete observations of reality.

There was a bag of gummy bears lying on the table. As the comedian performing this evening was from Scotland, I didn’t understand much of what was said. Therefore, I focused on the bag. It had all six colours of bear. How disappointing. I asked myself, ‘What is the probability of getting only one bag with the same colour?’ Was it even possible? Or does Haribo, the company that produced these sweets, secretly rig the packing process so that there will always be a variety of colours and flavours in a bag?

Therefore, I started a survey in 2018. Had I ever imagined how long it would go on for, or how many gummy bears I would end up counting, I would have stopped immediately. However, I was unaware of this, so I created a simple spreadsheet in Google Spreadsheets, shared the shortened link with every new lecture I gave, and counted the number of gummy bears, as well as collecting some demographic data about the students. From the beginning, one idea was good: I only collected some demographic data, such as body height, age, and the most popular flavour of gummy bear. Therefore, the data is from real people, but is anonymised as much as possible.

While you are reading these lines, the dataset is still growing. Currently, we have \(1134\) bags containing \(1.0365\times 10^{4}\) bears, counted by \(1031\) people. There are more bags of gummy bears than people who have counted them because some people counted the bears but did not enter any demographic information, or only entered it in part. Therefore, there is a lack of data. The beauty lies in the numbers: over one thousand bags and people, as well as over ten thousand bears counted in nearly ten years.

Therefore, we come to my urgent question. Given my average speed at counting bags during my lectures, is it possible that I will find a bag containing only one colour in my lifetime? Follow me through this chapter to find the answer!

4.2 Theoretical background

4.3 Data

I started generating data in 2028. From the beginning, the data generation process has been very hands-on and has not changed since. There is an empty Google spreadsheet with a shortened link. The process has not changed in any of the years that I have conducted the survey. First, I distribute the bags of gummy bears among the students. Each student gets one bag and can eat all the sweets after the survey. Then I present the link, and all the students join the same Google spreadsheet. Depending on the size of the lecture, it can get a little chaotic, but in the end, all the bears are counted and some demographic data is gathered.

The gummy bear data is split into two parts. The first part is a matrix showing the number of gummy bears of each of the six colours: dark red (d), light red (l), orange (o), yellow (y), green (g) and white (w). The second part shows the self-reported bear count and number of colours. Some errors can occur here. Because the bears and the number of colours have to be counted by the counter. Therefore, a discrepancy can occur between the colour matrix and the reported number of bears and colours. The second part includes some demographic information about the consumer. I wanted to know their age, height and favourite flavour of gummy bear.

4.3.1 Characteristics of gummy bears

Table 4.1: Extract of the gummy bears dataset of the characteristics of gummy bears. The year of the survey is reported, as well as the number of bears of each colour that were counted. The last two columns show the self-reported bear and colour counts.
Gummy bear color
Observed
Year Dark red Light red Orange Yellow Green White Bears Colors
2018 0 0 5 4 0 0 9 3
2018 0 3 1 4 1 1 10 5
2018 1 2 2 2 1 1 9 6
... ... ... ... ... ... ... ... ...
2026 0 1 2 3 1 2 8 5
2026 4 0 1 1 1 1 8 4
2026 1 0 2 2 4 1 10 5
Table 4.2: Extract of the gummy bear dataset of the self-reported counts (‘observed’) of the number of gummy bears, as well as their colours, and the deviation from the calculated counts from the gummy bear colour matrix (‘expected’) separated for the year and individual counting person. Each line represents one bag.
Number of bears (b)
Number of colors (c)
Year Observed Expected Deviation Observed Expected Deviation
2018 9 9 0 3 2 1
2018 10 10 0 5 5 0
2018 9 9 0 6 6 0
... ... ... ... ... ... ...
2026 8 9 -1 5 5 0
2026 8 8 0 4 5 -1
2026 10 10 0 5 5 0
Figure 4.1: Deviation between the sum of the gummy bear out of the count matrix and the self-reported counts of bears and colour in a bag. Most people deliver the same values; no deviation can be determined. The observed deviations scatter in both positive and negative directions. No bias can be observed. The following analysis uses the calculated and therefore expected values from the count matrix.
Table 4.3: Expected colour counts in a bag and the corresponding counts in the survey in absolute numbers as well as the percentage.
Expected colors (c) Number Percentage
1 0 0.00%
2 1 0.09%
3 84 7.41%
4 322 28.40%
5 534 47.09%
6 193 17.02%
Table 4.4: Expected bear counts in a bag and the corresponding counts in the survey in absolute numbers as well as the percentage.
Expected bears (b) Number Percentage
5 2 0.18%
6 3 0.26%
7 46 4.06%
8 365 32.19%
9 287 25.31%
10 300 26.46%
11 82 7.23%
12 33 2.91%
13 11 0.97%
14 4 0.35%
15 1 0.09%

4.3.2 Characteristics of the counters

Table 4.5: Extract of the gummy bear dataset of the most liked taste of the gummy bear, the gender, the grouped age and the body height of each counting person. The age groups have been created based on the continuous reported ages from each survey year. All measures are self reported.
Most liked taste Gender Grouped age Body height [cm]
lightred male 30+ 193
yellow female <22 159
white female <22 159
… … … …
darkred female 22-29 172
orange male 22-29 182
white male 22-29 175
Table 4.6: Information on the height, age and grouped age of the female and male gummy bear counters.
Characteristic female
N = 5281
male
N = 5151
height 169.38 (7.02) 183.93 (7.81)
    (Missing) 5 4
age 23.22 (5.64) 23.11 (4.42)
    (Missing) 2 4
age group

    <22 238 / 526 (45%) 206 / 511 (40%)
    22-29 247 / 526 (47%) 274 / 511 (54%)
    30+ 41 / 526 (7.8%) 31 / 511 (6.1%)
    (Missing) 2 4
1 Mean (SD); n / N (%)

4.4 Results

4.4.1 Gummy bears over the years of the survey

Figure 4.2: Dependency of the number of bears in a bag and the year of the survey. Due to the pandemic, data for 2020 is missing. The number of counted bags is shown above the histogram. The point and interval indicate the median and the range of quantiles. A clear shrinkage in the size of the gummy bear bags can be observed over the years.

4.4.2 Self reported body height

Figure 4.3: Information regarding the internal data integration of the individual reported height. No measuring tape was used to determine body size. The different heights for both genders are shown. (A) The top left shows body size as a histogram for both genders, with a bin size of 5 cm. (B) The top right shows density plots of body height for both genders. The line indicates the mean. There is a shift to the right for both genders. (C) The lower figure shows a beeswarm plot divided by gender. Men with a body height of 179 cm are especially rare. Men of this height have only started to appear in recent years.

4.4.3 Most liked gummy bear taste

Figure 4.4: Does Haribo know the common people’s preferences for gummy bears? Yes, but does Haribo care? Have they changed the production cycle to prioritise the most popular bears? (A) The number of people split by gender who liked which taste the most. (B) The number of gummy bears in all opened bags, split by colour. In the case of a uniform distribution, the same number of colours are produced, and each colour has an equal chance of being chosen. Therefore, the number of colours should be the dotted line.
Figure 4.5: Visualisation of the most liked colour taste of the gummy bears by a mosaic plot. (A) The plot on the left shows the separation of taste preferences by gender. The global chi-square test is not significant; there are no different taste preferences between the two genders. (B) The right plot shows the separation of the gummy bear taste preferences by three different age groups. The global chi-square test is not significant; there are no differences in taste preference between the different age groups of participants.

4.4.4 The gummy bear lottery

\[ f(y \mid\mu,\sigma^2)=\cfrac{1}{\sqrt{2\pi\sigma^2}} e^{-\cfrac{(y-\mu)^2}{2\sigma^2}}\quad -\infty<y<\infty \]

\[ \begin{align*} Pr(c=1|b) &= \tfrac{1}{6}^{b-1} \mbox{\; for\;} b \in \{1,..., 15\} \\ Pr(c=2|b) &= 14.53 \cdot \tfrac{1}{3}^{b} \mbox{\; for\;} b \in \{6,..., 15\}\\ Pr(c=3|b) &= 17.5 \cdot \tfrac{1}{2}^{b} \mbox{\; for\;} b \in \{6,..., 15\}\\ Pr(c=4|b) &= 2.98 \cdot 0.7607^b \mbox{\; for\;} b \in \{6,..., 15\}\\ Pr(c=5|b) &= 0.521 \cdot e^{-\tfrac{(b-10.02)^2}{(2\cdot3.39^2)}} \mbox{\; for\;} b \in \{6,..., 15\}\\ Pr(c=6|b) &= 0.00419 \cdot b^2-0.00303 \cdot b-0.12376 \mbox{\; for\;} b \in \{6,..., 15\}\\ \end{align*} \]

Figure 4.6: Probability of observation of the number of colours given the number of bears in a bag. The first nine bag sizes could be solved using combination and by counting the number of colours in a bag of a given size. However, the results for 10 to 12 bears in a bag had to be obtained by simulation and drawing 10 million bags due to performance issues. Gemini 3.5 (Flash-Lite), Excel and R with {nlme} were used to optimise the modelling and generation of the shaded graphs.
Figure 4.7: Percentage of observed numbers of colours in a bag given different bag sizes. On the left are the observed and expected percentages of eight bears in a bag. In the middle, the observed and expected numbers of colours in a bag containing nine bears. On the right, the observed and simulated percentages for a bag containing 10 bears. The simulation was necessary because calculating the expected values would take too long.
Table 4.7: The final table to judge the probability of observing c colours in a bag containing between 6 and 15 gummy bears. It can be assumed that bags containing fewer than six bears will never be produced. With seven gummy bears in a bag, the likelihood of finding only one colour is approximately 1 in 46,000.
Bears
Combined
Simulated
Modeled
Colors 6 7 8 9 10 11 12 13 14 15
1 0.01% 0.00% 0.00% 0.00% 0.00% 0.00% 0.00% 0.00% 0.00% 0.00%
2 1.99% 0.68% 0.23% 0.08% 0.03% 0.01% 0.00% 0.00% 0.00% 0.00%
3 23.15% 12.90% 6.90% 3.60% 1.85% 0.95% 0.48% 0.21% 0.11% 0.05%
4 50.15% 45.01% 36.46% 27.76% 20.29% 14.47% 10.11% 8.51% 6.47% 4.93%
5 23.15% 36.01% 45.01% 49.66% 50.65% 48.98% 45.63% 35.40% 26.15% 17.71%
6 1.54% 5.40% 11.40% 18.90% 27.18% 35.60% 43.78% 54.50% 65.51% 77.35%

With seven bears in a bag, I would like to have opened roughly 46,656 bags to find a bag with only one color of bear. With eight bears in a bag, it would be an astonishing 279,936 bags. Currently, I have an average of 185 bags counted per year among my lectures in the last five years. I still believe that the number of bears in a bag will remain stable at eight in the coming years. In this case, on average, I would need 1513 years before I would find one of a different colour. Even if Haribo reduces the number of bears in a bag to seven, it would still take me 252 years to reach my goal. Therefore, in answer to my initial research question, I realise that I have little chance of achieving my goal in my lifetime. But this will not stop me — maybe I’ll get lucky and the next bag will contain only one colour. The only assumption I have to make is that Haribo does not rig the filling process of the bags and removes any coloured bags that appear. It would be a bittersweet symphony.

The Table 4.8 shows a collection of odds for rare events that are not taken too seriously. Therefore, finding a one-coloured bag containing eight bears would be as unlikely as being bitten by a snake or hit by a comet in a lifetime. Neither event is more preferable than a bag of coloured bears. The table should not be taken too seriously. I gathered the information from a quick web search and did not consult any scientific references.

Table 4.8: A not-so-serious table of rare events and their odds of happening.
Event Odds
Odds to win Lotto 6 out of 49 \(1:15,537,573\)
Annual risk of being killed in a plane crash \(1:11,000,000\)
Lifetime odds to be hit by a comet \(1:1,600,000\)
Bag of 8 bears with one color \(\mathbf{1:279,936}\)
Bitten by a snake in everyday life \(1:40,000\)
Struck by lightning in a lifetime \(1:15,300\)
Lifetime odds of dying in a car crash \(1:100\)
Getting cancer in a lifetime \(1:5\)
Odds to be eaten by a polar bear \(>0\)

Well, dear reader, we’ve come to the end of this section. For me, the journey continues—I’ll keep tearing open packets and counting the bears. But you’re probably eager to learn more about the statistics by now. That’s what we’ll cover in the next few chapters.

4.5 Alternatives

Further tutorials and R packages on XXX

4.6 Glossary

term

what does it mean.

4.7 The meaning of “Models of Reality” in this chapter.

  • itemize with max. 5-6 words

4.8 Summary

References