Interactive Data Tools
Interactive demos for data science and data analysis concepts
Data fundamentals
Types of Data
Every variable in a dataset falls into one of four types, and knowing which one shapes what you can do with it. Nominal variables are categories with no natural order, like a dog's breed. Ordinal variables are categories with a natural order, like a satisfaction rating. Discrete numeric variables are counts — whole numbers you arrive at by counting. Continuous numeric variables are measurements that can take any value, down to as many decimal places as you can measure.
Join Types
A lot of real-world reporting errors — missing rows, doubled counts, totals that don't add up — trace back to merging two tables with the wrong join type. This chart makes that concrete: two small tea-club tables — members and their orders — get joined together. Pick a join type below and watch which rows survive, and where NULLs appear. A few other named join types exist (semi-joins, anti-joins, and the like) — they behave like the ones here with an extra filter layered on top, so they're left out here to keep things focused.
Statistics
Descriptive Statistics
A small set of data points sits on a number line. Drag them around, add or remove points, and turn on different descriptive statistics — mean, median, mode, minimum and maximum, standard deviation, interquartile range, and a boxplot — to see how each one responds as the data changes.
Sampling Methods
Getting a sample wrong can quietly produce wrong conclusions — whether you just grab whoever's easiest to reach, or assume any "random-looking" group of people is good enough. This demo shows why: a population of 60 people from three countries is shown, each with their own average daily tea consumption. Draw a sample using a convenience (biased) pick, simple random sampling, or stratified random sampling, and compare it to the true population. Stratified sampling is one method built to guard against this kind of error — several others exist too, each suited to different situations, but aren't covered here.
Regression & Correlation vs. Causation
Testing whether two things are correlated is one of the most useful tools in data analysis — it can surface a real relationship worth acting on, or show that something you assumed mattered actually doesn't. But it's also easy to be misled: a strong correlation can be coincidental, driven by a hidden third factor, or point the wrong way entirely once you split the data into subgroups. Six small interactive demos below cover both sides — two on how linear regression works, and four on ways correlation can mislead you if you take it at face value.
Data quality
Missing Data Imputation
Real-world data goes missing all the time — a sensor drops offline, a form field gets skipped, a survey section never reaches part of the sample, a record gets accidentally deleted. The ideal fix is often to collect more data, but that's not always possible: if a weather station's thermometer goes offline for an afternoon, there's no going back to re-measure that afternoon's temperature after the fact. Imputation is one way to handle a gap like that — instead of discarding the good data around it or leaving a hole, you fill it in with a plausible estimate based on the pattern the rest of the data shows. This chart demonstrates that with a scatter of tea steep time (minutes) vs. strength (1-10) — the true relationship is roughly linear, then curves as strength approaches its maximum. Choose a window of time where data goes missing, then click through different ways of filling in the gap.
Outliers
An outlier is a value that doesn't fit the pattern of the data around it — sometimes a one-off mistake, sometimes a sign something systematic is going on. The examples below are deliberately obvious, but real outliers usually aren't, which is why there are quantitative tests for how likely a point actually is one: a value beyond the 1.5× IQR whiskers, more than 2-3 standard deviations from the mean, unusually far from a fitted trend line, or — for skewed data where the mean and standard deviation are themselves thrown off by outliers — a modified z-score built from the median instead. This page has two examples of how a single unusual value can distort an analysis, and what changes once you flag it.
Binning: Numeric to Ordinal
Too many distinct values (or categories) can overcomplicate an analysis — the more of them there are, the harder real patterns are to see and the more fragile later steps become, an effect often called the curse of dimensionality. Reducing a fine-grained number down to a handful of ordered categories is one way to guard against that: a marketing campaign, for instance, doesn't need to know a customer's tea strength preference is exactly 6.4 out of 10 — "regular" is enough to act on. Each point below is one cup of tea's measured strength, on a scale from 1 to 10; here it gets converted into three ordered categories — "weak", "regular", and "strong". Where those category boundaries land is a real decision: drag the two cut lines yourself, or pick a method that computes them automatically.
Data visualization
Raw Visual Input
Before your brain groups anything into patterns, before it even decides what matters, your eyes have already run their own first pass: light lands on the retina, gets converted into signals for contrast, edges and colour, and gets routed toward the visual cortex — all before a single conscious thought happens. Two things about that first pass matter enormously for anyone building a chart. First, colour is a fragile signal on its own: poor contrast, or leaning on hue with nothing else backing it up, can distort what a reader sees before their brain even starts interpreting it. Second, not every visual system receives that signal the same way — roughly 1 in 12 men and 1 in 200 women have some form of colour vision deficiency, and plain low contrast is a problem for a great many more readers than that. The tabs below let you see effects like these live.
Before you can blink (Pre-attentive Processing)
Some visual differences register the instant you look. A single red dot in a field of grey ones seems to announce itself — you haven't searched for it, it is just there. This is pre-attentive processing: your visual system reads a few basic features, colour among them, everywhere at once, before focused attention gets involved. The catch is that it holds for only one feature at a time. Add a second and a third competing colour and the shortcut breaks down, and you are back to checking items one by one. This demo puts that to the test with a quick counting task: glance at a grid of digits, then say how many of one digit you saw.
Gestalt Principles
Before you consciously compare a single value on a chart, your brain has already grouped what it's looking at — into clusters, categories and stand-out elements — using a handful of built-in rules psychologists call the Gestalt principles. A chart that works with those rules feels effortless to read; one that fights them makes the reader do conscious work your design should have done for them. Seven principles are commonly grouped under this name: Simplicity, Continuity, Proximity, Similarity, Focal Point, Correspondence and Figure/Ground. Five get their own interactive tab below, each on the same small set of numbers so only one thing changes at a time. The other two show up differently: Simplicity is really the idea running underneath all five tabs rather than a demo of its own, and Correspondence — the risk of leaning on colours like red and green to mean "bad" and "good" — appears as a caution inside the Focal Point tab instead of a full tab of its own.
What Catches Your Eye First (Visual Search)
Some visual features jump out at you instantly, before you've consciously started looking — a red dot among blue ones, a tilted line among upright ones. Vision scientists call this pre-attentive processing: your visual system picks certain features out of a whole scene in parallel, in a fraction of a second, before any deliberate, one-by-one searching begins. The tell is reaction time: a genuinely pre-attentive feature stays roughly as fast to find no matter how big the grid gets. Combine two features so neither alone marks the target, though, and that shortcut disappears — you're back to checking items one by one, and reaction time reads reliably higher than any single feature's. Not every pairing loses the shortcut equally badly, either: some combinations stay closer to single-feature speed than others, though none of them gets it back completely.
Clear the Clutter (Attentive Cognition)
Attentive cognition is the last stage of turning a raw visual signal into something understood — actively looking, comparing and interpreting, rather than just noticing. Two things shape how well that works: how clearly a chart is designed in the first place, and how the mind organizes many individual items into a manageable few. The two tabs below each explore one of those: fixing a deliberately badly designed chart, and sorting a shared set of items into groups of your own choosing to see that second idea, chunking, do its work.
Machine learning & AI
ML Techniques in Miniature
Machine learning, in a nutshell: give a system some data to learn a pattern from, then apply what it learned to new data it hasn't seen before to make a decision. The four small, hands-on demos below each show a different piece of that process, grounded in a tea-brewing or tea-shop scenario — not a survey of the field, just a chance to see each technique actually run on data you can watch and control.
Market Basket Analysis (Data Mining)
Data mining is about finding patterns already sitting in data you've collected — not, like machine learning, using those patterns to predict something about a new case you haven't seen, and not, like AI, acting on that prediction. Market basket analysis is a classic example: it just looks for which items tend to get bought together. Leaf and Kettle is a tea shop whose point-of-sale system logs every order as a list of items bought together — a "basket" — and whose counter carries a small stand of pencils and paper for the study crowd alongside the tea. Below are 100 real-looking orders from Leaf and Kettle, each basket icon in the grid doubling as a percent chart — pick two items and see how often they actually turn up together.
AI or Augmented Intelligence?
Artificial Intelligence usually means a system that decides and acts entirely on its own; Augmented Intelligence means a system that assists a person, who still makes the final call — the same underlying model can be run either way, and which one you're building comes down to how much you trust its decisions, not how smart it is. A tea shop wants to know: can a simple computer program be trusted to sort incoming stock as Green or Black tea on its own, or should a person always double-check it? The program looks at a new sample's steep time and water temperature, compares it to cups it already knows the answer for, and makes a guess — plus a confidence score for how sure it is. Real systems (fraud detection, content moderation, medical-image checks) use exactly this kind of confidence cutoff to decide when to trust a model and when to send it to a person instead.
Probability & simulation
Base Rates and the "Positive Test" Paradox (Bayesian Statistics)
Bayesian statistics is about updating a belief in light of new evidence — starting from how likely something was beforehand (its "base rate") and combining that with how reliable the evidence actually is, rather than trusting the evidence in isolation. The classic place this trips people up is a positive test result: even a highly accurate test can be mostly wrong in practice if whatever it's testing for is rare enough to begin with. A tea supplier receives large shipments of tea, and a small fraction of any shipment turns out to be contaminated — a food-safety problem that has to be caught before it ships to customers. A quality-control test screens every batch, but no test is perfect: it can flag a contaminated batch correctly, flag a clean one wrongly, or miss real contamination and clear it as fine. This chart simulates a full 1,000-batch shipment against your own choice of how rare contamination is and how accurate the test is, showing what actually happens to every batch.
Game Theory
Game theory studies how people make decisions when the outcome depends on more than just their own choice — sometimes because someone else is choosing too, and sometimes because someone else simply knows something you don't. Two classic puzzles below: a game-show scenario where switching your guess is secretly the better strategy, and a pricing standoff where the individually safest choice for each side leaves both worse off.
Pi with Teabags (Monte Carlo)
Monte Carlo methods estimate an answer by running a large number of random trials and looking at the aggregate outcome, rather than solving the problem directly — useful whenever an exact calculation is too complex, or there's no clean formula for what you're after in the first place. The more trials you run, the closer the estimate typically gets to the true value. Someone's pulling a box of tea bags off a shelf in a tea shop's stockroom, and the box slips, scattering tea bags across the floor. The tile underneath happens to have a circle inscribed inside its square border. Since where each bag lands is effectively random, counting how many land inside the circle versus outside it turns out to be a genuine way to estimate the number pi — a circle with radius 0.5 fits exactly inside a 1x1 square and covers pi/4 of its area, so four times the fraction landing inside is an estimate of pi itself.
The Tea Urn (Process Modelling)
Process modelling is a way of tracking how something changes over time — the same basic idea applies whether it's water in a tank, money in a bank account, energy in a battery, or cups of tea in an urn. Every version comes down to two things: a stock (however much of something you currently have) and the flows adding to it or draining it. It's useful for understanding how a process behaves over time, for planning ahead when you can influence some levers (like when to refill) but not others (like demand), and for seeing how a delay between a decision and its effect can throw off even a reasonable plan. You're running the tea counter at a busy café. A 20-cup urn has to stay stocked through opening hours, but customer demand follows a rough rush-hour pattern that's never quite the same two days running — busier some mornings, quieter some lunches, always with a bit of a surprise.