Skip to content
BusinessTech AcademyBusinessTech AcademyData Dive

Why this is here

BusinessTech Academy is not a mathematics academy, and this is not a statistics course. It is a side piece: half an hour at a time, a chart on the screen, and three questions. This page says why that is worth doing and where the reasoning comes from, because a side piece with no stated reason is the first thing cut.

The field's own position: not in math class

In April 2024 five national bodies — the National Council of Teachers of Mathematics, the National Science Teaching Association, the American Statistical Association, the National Council for the Social Studies, and the Computer Science Teachers Association — signed one sentence together:

Data science bridges disciplines and thus should be introduced and taught across the curriculum in K-12 schools to help develop informed users of data.

Across the curriculum. Not inside mathematics. They name the reason plainly: a gap exists between the concepts taught in math and the data skills needed by other disciplines.

Their third principle is this project's whole editorial position:

Making sense of and understanding the data must be the focus rather than the mechanics of computations and calculations.

That is the licence to run a data strand inside a business and technology academy without becoming a second math class, and it is signed by the math teachers' own association.

What a BusinessTech student gets out of it

Steve Levitt, writing the preface to the guidelines this project takes its light from:

Basic data fluency is a requirement not just for most good jobs, but also for navigating life more generally, whether it is in terms of financial literacy, making good choices about our own health, or knowing who and what to believe.

Good jobs. Money. Health. Knowing who to believe. For an academy that puts students into business and technology, that is the argument entire, and it needs no chapter on the normal distribution to make it.

GAISE II is the light, not the syllabus

The Pre-K–12 Guidelines for Assessment and Instruction in Statistics Education II is where the shapes on this site are grounded, and the small gold citations on the library cards point into it. It is a serious document and it is honest about what it is: its own preface records that the first edition was written to enhance the statistics standards inside NCTM's Principles and Standards for School Mathematics. Levels A, B and C, a four-step investigative process, multivariate thinking, probabilistic reasoning.

We take its reason and its data sets. We do not take its apparatus. Nothing on this site asks a student what level they are at, and the four sections of the shape library are not GAISE II levels — they sort by visual complexity so the page can be walked.

What we actually do: provenance

The joint statement's fourth principle is the one this site was built around before we had a name for it:

Students must learn to question the sources of data... The provenance of data is fundamental to understanding data quality; we must be honest in how the data were collected and assumptions that were made during collection.

The same move is the centre of the Level C column in the Harvard Data Science Review: the authors argue for dedicating ample classroom time to curating, questioning the provenance (documentation of the source and creation process) of data; that is, truly understanding data origins and study design.

That is what every footer on every file here is for. Each CSV ends with its publisher, its URL, its licence, the query that produced it and the date it was fetched. Every chart names its source beside it. A file named anywhere on this site is a link to that file. None of that is housekeeping. It is the fourth principle, in the only form a fifteen-year-old can check for themselves.

Skeptics, not cynics

The word the HDSR authors use for the goal is healthy skeptics — students asking questions about the process that brought the data behind a visualization into existence:

Where and what data were utilized to create this visualization? · Why was this visualization created and by whom? · Who collected the data and why? · Who funded the data collection?

Every one of those is a question about origin. None of them is a verdict. That distinction is the whole difficulty of the dives, and the teacher note in the dive bank says so: given permission to be suspicious, a fifteen-year-old will call everything a lie by chart four, and that is not skepticism — it is a different way of not thinking. Skepticism is a question about provenance. Cynicism is a verdict without one.

The eight questions

What a class can ask of any chart it meets, from the HDSR column. They are the long form of the dives' first question, and they work on a chart nobody here made:

  1. What was the purpose for collecting the data?
  2. Who collected the data?
  3. Who funded the data collection and research?
  4. Was the data collected using an observational study or an experiment?
  5. Whom was the data collected from?
  6. When was the data collected?
  7. Where was the data collected?
  8. How many people participated?

The three questions every dive runs — what does this shape leave out, what does the chart actually say, and where does it touch your own street — are the short form, cut to fit twenty-five minutes.

Sources

The HDSR column's own worked example uses the American Time Use Survey, which ships here as gaise2_atus_time_use_sample.csv.