PDF regeneration pending (Noah)

# CourseTrees Data Paradigm

Last updated: July 9, 2026

CourseTrees is independent and student-built, not affiliated with or endorsed
by any university. It is free to explore, and optional paid plans may be offered in the future; CourseTrees will never sell data or paywall public catalog data.
We can only ask students to trust the map if we are careful with data. This
document explains, in plain language, what data we hold, where it comes from,
and how we treat it under the privacy and ethics frameworks that apply to us.

In short: the course data we display is factual and public, the grade
distributions we show are aggregate and de-identified, we suppress any group
small enough to risk identifying a student, we never sell personal data, and
we honor access, deletion, opt-out, and takedown requests.

## What data we hold

- **Public catalog data:** course codes, titles, descriptions, prerequisites,
  and offering history, gathered from each school's public course catalog.
  This is factual information about courses, not about people.
- **Aggregate grade distributions:** where a school publishes them or makes
  them available through public-records channels, we show the distribution of
  grades for a section or term as counts and percentages. We also show
  student-reported grade distributions from course-feedback surveys when
  available, labeled "Student-reported." These are group statistics. They are
  not tied to any student.
- **Public course and professor ratings:** coarse, aggregate signals drawn
  from public sources, shown under a general community label, never attributed
  to a named individual reviewer.
- **Optional account data:** if you sign in with a school email to post, we
  store your email, a small profile, and what you post. Browsing needs no
  account, and your school, courses taken, and plans stay in your own browser.
  Full detail is in our [privacy notice](/privacy).

## FERPA: why our data is not protected education records

The Family Educational Rights and Privacy Act (FERPA) protects a student's
education records: personally identifiable records that a school or its agents
keep about that individual student, such as their transcript, grades, and
enrollment. It governs institutions that receive federal funding and the
records they hold about identified students.

CourseTrees is not a school, and we do not hold education records. The two
kinds of data we display are different in kind from a protected record:

- **Public catalog data** describes courses, not students. A course
  description or a prerequisite list is not personally identifiable
  information about any individual.
- **Aggregate grade distributions** describe a group, not a person. A
  statement that a section of 120 students earned 30% A grades says nothing
  identifiable about any one of them. FERPA itself treats properly
  de-identified, aggregated information as outside the definition of a
  protected record, as long as a student cannot reasonably be identified from
  it.

The key safeguard is de-identification. To keep aggregate data from becoming
re-identifying, we apply the minimum cell-size floor described next.

## Our grade-suppression floor: a minimum cell size of 5

We never display a grade distribution for a group small enough to risk
re-identifying a student. Our floor is a **minimum cell size of 5**. If the
number of students behind a distribution, or behind a cell within it, falls
below 5, we suppress it rather than show it.

Small groups are where aggregate data stops being anonymous. In a section of
three students, a single grade can point to a single person, especially when
someone already knows who was enrolled. A floor of 5 is a widely used
threshold for reporting aggregate education and health statistics without
exposing individuals. We treat it as a hard minimum, not a goal.

## GDPR: rights for people in the EU and EEA

The General Data Protection Regulation (GDPR) governs the processing of
personal data of people in the European Union and European Economic Area. Most
of what we show is not personal data at all, but the account data of a
signed-in user is, and we treat it accordingly. If GDPR applies to you, you
have the right to access the data we hold about you, to have wrong data
corrected, to have your data erased, to restrict or object to certain
processing, and to receive your data in a portable format.

We rely on your consent, given when you choose to sign in and post, and on our
legitimate interest in running a study aid, as our bases for processing the
small amount of account data we hold. We do not sell personal data and we do
not use it for advertising. To exercise any right, email
coursetrees@gmail.com.

## CCPA and CPRA: rights for California residents

The California Consumer Privacy Act (CCPA), as amended by the California
Privacy Rights Act (CPRA), gives California residents rights over their
personal information. If you are covered, you have the right to know what
personal information we collect and how we use it, to delete the personal
information we hold about you, to correct inaccurate information, and to opt
out of the sale or sharing of personal information.

We do not sell or share personal information as those terms are defined by the
CCPA and CPRA. There is nothing to opt out of on that front, and we will not
start selling your data. To make a request, email coursetrees@gmail.com. We
will not treat you differently for exercising any right.

## Data minimization

We collect as little as the site needs to work. Browsing requires no account.
Your school, the courses you mark as taken, and any plan you build stay in your
own browser and are never uploaded to us. Analytics are coarse and
non-identifying: we strip course, plan, and other student-specific values from
the URL before recording a page path, we do not send raw search text or course
titles, and we do not run session recording. If you never sign in, we hold no
account for you.

## How Cam uses AI

Cam uses third-party inference, which consumes data-center energy and can consume water through cooling and electricity generation. The footprint varies materially by model, hardware, utilization, data-center location, grid mix, cooling system, prompt length, output length, and cache behavior.

CourseTrees has not directly measured Cam's per-answer energy, carbon, or water use and will not present another system's number as Cam's.

To keep Cam bounded, CourseTrees uses these controls:

- small-model-first routing
- verified-answer caching
- bounded tool/token loops
- beta restriction
- daily limits
- paid fallback routing disabled by default

The references below are the sources we used for pricing, public rate-limit
context, and measurement caveats. They explain why footprint claims vary
materially by model, hardware, utilization, location, grid, cooling, prompt
and output length, and cache behavior.

- [Models & pricing](https://platform.claude.com/docs/en/about-claude/pricing) — Anthropic, verified 2026-07-09. Current Claude Haiku 4.5 base pricing used for the conservative paid ceiling.
- [Rate Limits](https://console.groq.com/docs/rate-limits) — Groq, verified 2026-07-09. Published base Free-plan limits are a public reference, not proof of CourseTrees' current tier or exact limits.
- [Power Hungry Processing: Watts Driving the Cost of AI Deployment?](https://arxiv.org/abs/2311.16863) — Luccioni, Jernite, and Strubell, verified 2026-07-09. Shows that inference energy varies substantially by task, model, hardware, utilization, and deployment context.
- [Making AI Less "Thirsty"](https://arxiv.org/abs/2304.03271) — Li, Yang, Islam, and Ren, verified 2026-07-09. Explains that water impact depends on location, grid mix, cooling system, timing, and measurement method.
- [Measuring the environmental impact of delivering AI at Google Scale](https://arxiv.org/abs/2508.15734) — Elsworth et al., verified 2026-07-09. Demonstrates full-stack production measurement; its provider-specific median is not a Cam estimate.

## How we treat catalog sources: robots and Terms

We gather public catalog data with automated tools, at a respectful rate that
does not burden a school's systems, and we attribute each source with a link
back to the official catalog. Before we add a source, we review its published
Terms and its robots rules, and we design our collection to stay within normal,
respectful use rather than aggressive bulk harvesting. When a source signals
that automated access is not welcome, we respect that signal. Coverage varies
by school as a result, and we would rather have less data than data we should
not have taken.

## The ACM Code of Ethics

We hold ourselves to the professional standards in the Association for
Computing Machinery (ACM) Code of Ethics and Professional Conduct. In practice
that means:

- **Avoid harm.** We weigh the effect of our data choices on students, and we
  suppress anything that could identify an individual.
- **Respect privacy.** We collect the minimum, keep personal data out of what
  we publish, and protect the little account data we hold.
- **Be honest and trustworthy.** We describe our data and its limits plainly,
  including where coverage is thin or a source is only the catalog.
- **Respect the work and rules of others.** We attribute sources and follow
  their Terms and robots rules.

## Opt-out and takedown

We provide a clear way to have data changed or removed:

- **Instructor or individual opt-out.** If you are an instructor or an
  individual and you want data associated with you removed or excluded, email
  coursetrees@gmail.com and we will act on it.
- **Per-source takedown.** If you represent a university or a data source and
  you want data corrected or removed, email coursetrees@gmail.com. We will
  respond promptly and, where appropriate, stop collecting from that source.
- **Correcting a specific course.** If a course detail looks wrong, use the
  report button in any course panel, or email us.

## Contact

Questions, opt-out requests, and takedown requests all go to the same inbox:
coursetrees@gmail.com.
