Treco

Free, open-source, privacy-first AI usage tracking browser extension.

View the Project on GitHub medaharrat/treco

Environmental methodology

This document is the source-of-truth companion to the in-app Settings → Environmental methodology page (packages/extension/src/ui/pages/MethodologyPage.tsx), which renders the same profile database live from packages/core/src/environment/modelProfiles.ts.

The calculation, step by step

tokens (input, output)
   │  × per-token or per-request energy coefficient (IT-equipment energy)
   ▼
IT-equipment energy (Wh)
   │  × PUE (Power Usage Effectiveness)
   ▼
Facility energy (Wh)
   │  × grid carbon intensity (gCO2e/kWh)         │  × WUE (mL water/kWh)
   ▼                                                ▼
Estimated CO2e (g)                             Estimated water (mL)

Implemented in packages/core/src/environment/calculate.ts. Every step is a documented, overridable assumption, not a hidden constant.

Token estimation

Handled by packages/core/src/tokens/estimate.ts. In the near-total absence of providers exposing real tokenizer counts to a browser extension (see docs/LIMITATIONS.md), token counts are estimated by blending two well-known rules of thumb for English text:

The two estimates are averaged for a central value, with an explicit uncertainty range (roughly 0.75×–1.35× the central estimate) to reflect that different tokenizers, non-English text, code, and markdown all diverge from this heuristic by a meaningful margin. Every token count carries a method tag:

Model profile database

packages/core/src/environment/modelProfiles.ts holds two kinds of entries:

1. Disclosed provider figures (confidence: "medium", never "high")

OpenAI and Google have each published approximate, aggregate per-query energy/water/carbon figures for their consumer chat products. These are the closest thing to “model-specific data” realistically obtainable, so they’re used directly as whPerRequest (a flat per-interaction energy figure) rather than a per-token coefficient, since that’s the granularity the providers themselves disclosed. Confidence is capped at "medium" because neither provider has published a full methodology — these are company-reported aggregates, not independently verified measurements.

These profiles are deliberately not tied to a specific named model version (e.g. “GPT-4o”). Providers ship new flagship models faster than any browser extension can track, and hardcoding a version number just guarantees the label goes stale within months. Instead, each disclosed profile describes “whichever model a provider’s disclosure was actually measured against at the time,” is stamped with the date it was last reviewed, and is expected to be updated whenever a provider publishes a new figure — not whenever they ship a new model name.

Extended-reasoning variants (models whose name suggests hidden chain-of-thought generation, e.g. “o1”/”o3”-class models) and multi-step “deep research” or agent/tool-use modes are deliberately excluded from these disclosed profiles and routed to their own fallback tiers instead — see resolveProfile.ts. A provider’s published average for its default chat model does not represent an interaction that may run many times more hidden reasoning or tool-call work per visible response.

2. Fallback tiers (confidence: "low", always isFallback: true)

Used for every provider/model without a credible public disclosure — which, honestly, is most of them. Models and modes are sorted into one of five broad tiers by keyword matching on the model/mode name detected from the page:

Each tier has its own per-token energy coefficients (output always weighted higher than input, since generation dominates inference cost), a wide uncertainty band (0.4×–2.5×, widest for the agentic tier in practice since its central estimate is itself the least certain), and an explicit note explaining what it’s applied to and why. This is intentionally never a single number: see packages/core/tests/environment.test.ts for tests asserting no fallback profile is ever presented at "high" confidence, that output coefficients are never lower than input coefficients, and that agentic-mode labels are never routed to a disclosed provider profile.

This keyword-matching approach is inherently a coarse, best-effort heuristic based on whatever model/mode label a provider’s page happens to expose - not a measurement of what actually ran. It is the largest source of uncertainty in the whole calculation, most of all for reasoning and agentic modes.

Other providers we checked

Only OpenAI and Google have a disclosed profile. Before adding one for any other provider, we checked what’s actually been published, rather than assuming either “no one else discloses anything” or reaching for a third-party estimate and treating it as if the company said it:

If any of this changes, the affected provider gets a real disclosed profile the same way OpenAI’s and Google’s did: built from their own published figure, not an estimate of one.

Global default assumptions

When no model-specific figure applies, these industry-average-order-of-magnitude defaults are used (all overridable per-profile, and the carbon intensity is user-overridable in Settings):

Assumption Default Basis
PUE (Power Usage Effectiveness) 1.56 Uptime Institute Global Data Center Survey 2024 industry-average PUE
Grid carbon intensity 442 g CO2e/kWh Order-of-magnitude global average electricity grid intensity (IEA electricity data)
WUE (Water Usage Effectiveness) 1800 mL/kWh Order-of-magnitude data center site water usage average

These are broad averages, not measurements of any specific facility. A user who knows their local grid’s carbon intensity can override it in Settings → Environmental methodology, which makes CO2e estimates more relevant to where they actually live.

Everyday comparisons

packages/core/src/environment/comparisons.ts converts energy figures into comparisons like “roughly as much electricity as running a laptop for X hours.” The constants behind these (a ~12 Wh smartphone battery, a ~50 W laptop draw, a ~9 W LED bulb) are shown as documented, order-of-magnitude assumptions — never as a precise conversion factor — and the UI always states explicitly that the comparison is approximate.

Sources & further reading

These are the same citations shown (with links) on the in-app Methodology page, so anyone can check the order of magnitude for themselves rather than taking this project’s word for it:

What this deliberately does not do