Back to Data Room
Methodology

How we read the evidence.

No scores. No fear. No affiliate links.

The Ingredient Lab translates peer-reviewed research and official regulatory assessments into honest, context-aware analysis of personal care ingredients for New Zealand and Australia. This page sets out exactly how that analysis is produced: what counts as evidence, how it is weighed, where the limits are, and why we will never reduce an ingredient to a single number. Our method is as public as our results. If it changes, we say why.

Version 1.0 Reviewed June 2026 · primary and regulatory sources only · human-reviewed before publication
01The principle

We do not score products.

Every other ingredient app gives you a number or a grade. We do not, and we never will. A score is a verdict that hides its own reasoning. It collapses three things that cannot be collapsed: how much of an ingredient is actually present, how strong the evidence behind a concern is, and how much that concern applies to you in particular.

Two ingredients can share a grade for opposite reasons, and a number cannot tell you which. It can only tell you to be afraid or to relax, and it asks you to trust it without showing its work. We take the opposite position. We show you the evidence, the strength of that evidence, the regulatory position, and the context, and you decide.

An informed reader does not need a score. A score is what you give a reader you are not willing to inform.

02Sources

The evidence we use.

A finding appears on this platform only if it traces to a primary source: a peer-reviewed study with a verified DOI, or a published assessment from a recognised regulatory body. No paper, no flag.

We do not build safety claims on ingredient aggregators or marketing databases. Those serve one purpose only, confirming chemical identity, which substance and which CAS number, never as evidence that an ingredient is safe or harmful. The regulatory assessments we anchor to include the Cosmetic Ingredient Review, the EU Scientific Committee on Consumer Safety, the US FDA, the European Chemicals Agency, Health Canada, AICIS in Australia, and the New Zealand EPA. Where these bodies disagree, we report the disagreement rather than quietly picking a winner.

03Weighting

How we weigh evidence.

Not all evidence carries the same weight, and treating it as if it does is how fear gets manufactured. We rank evidence in a fixed order: human population studies first, then human clinical studies, then whole-animal studies, then studies on isolated cells.

1
Human population studies
Strongest. What actually happens in people, at real exposures.
2
Human clinical studies
Controlled human exposure. Strong, but narrower in scope.
3
Whole-animal studies
A living system, but not a human one. Useful, with caveats.
4
Isolated-cell studies
Weakest. A real result, at doses and conditions skin may never meet.

A result in a dish is a real result, but it is the weakest kind. For this reason, a concern supported only by cell-level studies is capped at our lowest concern level for cancer, hormone, and organ effects. We will report that the signal exists. We will not inflate it into a risk the human evidence does not support.

We also report each finding at its own level and stop there. Where the science leaves a question open, for instance whether an effect seen in cells is reachable in the body, we say it is unresolved and present both sides. We do not adjudicate questions the researchers themselves have not closed.

04The two axes

Concern and confidence are different questions.

Most safety claims confuse two separate things: how worried should I be, and how settled is the science. We keep them apart, because they move independently.

Concern
How strong is the evidence that this ingredient causes harm. Driven by the strongest available evidence, and capped by its tier.
Confidence
How firmly the question has been answered. A settled question can carry high confidence even when the concern is low.

We can say with high confidence that no causal link has been established between a given ingredient and a disease, and that is a strong, defensible statement. What we will never say is that an ingredient is proven safe. Proving a negative to that standard is not how science works, and any platform that claims it is selling you a certainty it does not have.

Our verdicts are bounded for exactly this reason. "No causal link established at typical use levels" is a claim we can stand behind. "Safe" is not a claim anyone can stand behind.

05Dose and context

Context, not concentration.

A hazard is not a risk until you know the dose. Salt, water, and oxygen are all hazardous at the wrong dose. What matters for a real product is how much of an ingredient is present, in what kind of product, applied how often.

We assess each flagged ingredient against the regulatory limit for that product type, and against the total daily exposure a person gets from using several products at once, something single-product safety limits were never designed to capture.

This is only possible when a brand discloses its full ingredient list. Concentration is not published directly. The ingredient list gives order, not amount: ingredients appear in descending order down to one percent, below which they may appear in any sequence. We use that order as a rough position, never as a precise figure, and we tell you on every product whether full disclosure was available. When it was not, context is limited or unavailable, and we say so plainly.

This is a constraint we state, not a gap we hide. We retired the phrase "concentration-aware" precisely because it implied a precision the available data does not support. The honest word is context.

06The process

How a verdict is produced.

Behind every verdict is the same fixed process, run in the same order. The order is what makes a verdict auditable rather than an opinion. The principle that holds it together is simple: the AI extracts, the rules decide, and a human checks the result.

  1. 1
    Identity
    Every ingredient is resolved to a single canonical substance and its registered identifiers before anything else happens, so research about one chemical is never misattributed to a similarly named one.
  2. 2
    Retrieval
    We gather the primary literature for that substance, starting from the authoritative regulatory assessment and following its citations outward, rather than casting a wide keyword net and fighting the noise.
  3. 3
    ExtractionAI, grounded
    An AI system reads the literature and pulls out the relevant findings. Its only job is to locate and quote the exact text of the source, word for word. It does not score, judge, or decide. Every finding is tied to the verbatim sentence it came from, and if the system cannot ground a claim in the source text it stops and reports rather than filling the gap. This is enforced in the code, not left to good intentions.
  4. 4
    DecisionFixed rules
    Written, deterministic rules, the evidence hierarchy and the concern-and-confidence model set out above, are applied to the extracted findings to produce a draft verdict. The same inputs always produce the same output. This is the line that makes our analysis defensible: the AI finds the evidence, the rules decide what it means, and the two never blur.
  5. 5
    Human review
    A scientist checks every draft against the primary sources before it is published. The human gate is not removed for speed.
  6. 6
    Verification before release
    A verdict is released only once it passes against an independently authored reference for that ingredient, written from the primary sources without reference to the engine's output. Grading a system against its own work proves nothing.

We are calibrating this process now, which is why verdicts are published deliberately rather than at volume. A wrong verdict published quickly is worse than a correct one published slowly. The pause is the standard working, not a gap in it.

07What we assess

The eight categories.

We assess every product across eight categories. Each one is a plain-language question, not a technical label.

Skin & Eye Irritation
Will it irritate skin or eyes, and through what mechanism.
Hormone Disruption
Is there credible evidence of endocrine activity at the levels actually used.
Pregnancy & Fertility
Is there reason for caution in pregnancy or for fertility, on current evidence.
Cancer
Is it classified as a carcinogen by any recognised body, and at what exposure.
Organ & Systemic Toxicity
Can it affect organs or the wider body once absorbed.
Fragrance & Allergens
Does it contain declared allergens, and is the scent mixture disclosed.
Environment
What happens to it after it goes down the drain.
Data Transparency
How complete and verified is the information the brand actually provided.
08The labels

Reading a verdict.

Two labels carry most of the meaning on a product page. Read them together: one tells you the finding, the other tells you how much weight to give it.

Category status
ClearNo ingredient in this category meets our flagging threshold on the available evidence.
FlaggedSomething here is worth knowing about. A flag is information, not an instruction to avoid. Read the context.
Flagged · DisclosedFlagged, but the brand voluntarily disclosed more than the law requires. The disclosure is the positive.
Not assessableWe cannot reach a conclusion, almost always because the brand did not disclose enough to assess.
Evidence quality
Strong Multiple human studies, or a settled regulatory assessment, pointing the same way.
Moderate Limited human data, or consistent whole-animal evidence.
Weak A single study, animal-only, or cell-level evidence. Real, but the weakest kind.
No data No research was found. This is not the same as safe.
Not assessable The brand did not disclose enough for any conclusion to be reached.

When we find nothing, we say nothing was found.

We never translate a blank into "safe." Absence of evidence is not evidence of absence. A substance no one has studied is not a substance that has been cleared. It is a substance no one has studied, and you deserve to know the difference.

09Independence

Nothing we conclude is for sale.

Our analysis is funded by the people who read it, through subscriptions, and by nothing else. No brand partnerships. No affiliate links. No advertising. No paid placement. No verification badges sold to brands.

This is structural, not a promise of good behaviour. No revenue stream is permitted to influence a published verdict, because the moment a platform's income depends on what it concludes, its conclusions are compromised, whether or not anyone intends it. If a brand does well in our analysis, it is because the evidence put it there.

10Authorship

Who writes this, and how it changes.

The analysis is built by a scientist with direct experience manufacturing cosmetic raw ingredients, training in analytical chemistry, and doctoral-level work in research and scientific literature interpretation. Not a formulation chemist, but someone who understands how these ingredients behave, how they are characterised, and how to read a study for what it does and does not show.

This methodology is a living document. When our rules change, we record what changed and why. The standard we hold ourselves to is the one we ask you to hold us to: show the evidence, show its limits, and never pretend to a certainty the science has not earned.