A sustainability manager opens ChatGPT, pastes a list of expenses, and asks for a carbon footprint. The answer arrives in seconds, well written, with numbers and percentages. The problem only surfaces when an auditor, a client, or a regulator asks for the source of each emission factor used.

An LLM (large language model) such as ChatGPT, Claude or Copilot excels at writing, rephrasing or explaining a regulatory concept. It was not built to produce a defensible carbon figure, meaning a result a company can stand behind in front of an auditor, a client, or an authority. This article details what separates the two uses, and what a carbon management system placed under the AI layer adds that the AI alone cannot.

Table of contents

  • What an LLM is actually good at
  • Why an LLM alone does not produce a reliable carbon figure
  • What a defensible carbon figure requires
  • The right architecture: AI on top of a data system
  • How Kabaun structures this architecture
  • FAQ
  • Sources
  • What an LLM is actually good at

    An LLM is a statistical model trained to predict the most likely continuation of a text. That capability makes it genuinely useful for several tasks linked to a carbon footprint:

  • Writing and rephrasing: summarizing a methodology note, translating a CSRD requirement into plain language for a leadership team.
  • Explaining a regulatory text: summarizing an ESRS E1 standard, comparing two scope definitions.
  • Extracting unstructured text: spotting a piece of information in a supplier contract or a scanned invoice, provided a downstream system structures and validates that extraction.
  • Generating prompts and templates: drafting an audit checklist or the questions to ask a supplier.
  • These uses are covered in detail in our guides on AI prompts for CSR managers and AI prompts for a CSRD report. What these uses share: the LLM assists a language task. Calculating a carbon footprint is a data task.

    Why an LLM alone does not produce a reliable carbon figure

    The hallucination risk is not an anecdote

    A hallucination is a statement produced by an LLM that sounds plausible but is not true or verifiable. According to a study published by OpenAI researchers in September 2025, hallucinations partly stem from how models are trained and evaluated: a model that guesses a plausible answer statistically scores better than a model that answers "I don't know," which encourages confident but ungrounded answers.

    Applied to a carbon footprint, this mechanism is direct: ask an LLM for the emission factor of a specific item (a mode of transport, a material, an industrial process), and it can produce a plausible, well-formatted number with no verifiable link to an official database. Nothing in the answer lets you distinguish a real factor from a hallucinated one.

    No structured or historized activity data

    A carbon footprint is built from activity data: liters of fuel consumed, kWh purchased, kilometers traveled, tons of material bought. A general-purpose LLM has no structured memory of this data: it processes the text pasted into the conversation, with no persistent database linking each data point to its source (invoice, meter reading, supplier declaration), and no history carried from one reporting year to the next.

    Emission factors with no source or version tracking

    The GHG Protocol Corporate Standard, the international methodological reference for corporate carbon accounting, relies on converting activity data into emissions through a documented emission factor. A credible factor database (the ADEME Base Carbone, DEFRA, EPA) ties each factor to a source, a methodology, and an update date. An LLM queried without a verified connection to such a database cannot guarantee the source, the version, or the validity date of the factor it returns.

    A calculation that is not deterministic or reproducible

    An LLM generates its answers probabilistically: asking the same question twice can produce two different phrasings, and sometimes two slightly different results. A defensible carbon footprint must instead be reproducible: the same activity data and the same emission factor must produce exactly the same result, year after year, to allow comparison over time and verification by a third party.

    No multi-entity consolidation or scope management

    A group with several subsidiaries, sites or legal entities must consolidate its data under a defined boundary (operational control, financial control, or equity share, under the GHG Protocol). That consolidation requires a data structure organized by entity and site, with consistent aggregation rules. An LLM does not maintain that structure: every conversation starts from a text, not from a data model.

    No audit trail

    An auditor or a verifier reviewing a carbon footprint traces every number back to its source: who entered the data, when, with what supporting document, which factor was applied and who validated it. A conversation with an LLM produces no timestamped, tamper-evident audit trail of that kind.

    What a defensible carbon figure requires

    A carbon figure becomes defensible, meaning it can withstand third-party scrutiny, when it meets requirements the conversational format of an LLM does not natively cover:

  • Traceability: a documented source for each data point. An LLM alone has no structured memory linked to a supporting document.
  • Sourced factors: an official, dated, versioned factor database. An LLM alone has no verified connection to a public database.
  • Reproducibility: same data and same factor, same result. An LLM generates probabilistically, with no guarantee of an identical result.
  • Consolidation boundary: operational or financial control rules per entity. An LLM alone has no multi-entity data model.
  • Audit trail: a timestamped history of entries and validations. A conversation leaves no persistent record of this kind.
  • Human control: explicit validation before publication. An LLM alone has no built-in validation mechanism.
  • These requirements are not specific to CSRD: they flow directly from the GHG Protocol methodology, described in our article on carbon accounting, and from the logic of sourced emission factors.

    The right architecture: AI on top of a data system

    The mistake is not using AI for carbon accounting, it is asking AI to replace the system instead of operating on it. A reliable architecture places the LLM on top of a structured system, rather than in place of that system: some vendors in the sector describe this layer of structured, historized and sourced data as the foundation, or "operating system," on which AI runs, by analogy with a computer's operating system organizing resources beneath the applications.

    Concretely, that means:

  • A persistent activity data layer, where every line is linked to a source and historized by reporting period.
  • A sourced and versioned emission factor database, updated and documented, rather than a number generated on the fly.
  • A deterministic calculation engine, compliant with the GHG Protocol, which always applies the same rule to the same dataset.
  • An AI layer that acts through tools, on this structured data, rather than on text pasted into a chat window.
  • Explicit human validation before any action or figure proposed by the AI enters the final footprint.
  • This agentic logic, where AI acts through audited tools rather than producing an isolated answer, is detailed in our article on the AI carbon agent. It also matches the selection criteria developed in our guide on AI-powered CSR software: what makes a tool reliable is not the presence of a chatbot, but exactly where the AI sits in the data chain. Once that architecture is in place, automation becomes relevant category by category, as detailed in our article on automating carbon accounting.

    How Kabaun structures this architecture

    Kabaun is a carbon management platform built on this logic: data first, AI on top, human validation before every structuring action.

  • A calculation engine compliant with the GHG Protocol Corporate Standard (product code CBC-001), using sourced emission factors from referenced public databases (ADEME, DEFRA, EPA, among others), rather than a number generated by a language model.
  • Multi-entity and multi-site management with group-level consolidation (GDD-004), to cover an organization's actual boundary.
  • Klem, Kabaun's AI assistant, answers natural-language questions on the footprint data (IA-004) and acts as an agent on specific tasks: invoice extraction, matching accounting lines, flagging import anomalies, simulating reduction trajectories (IA-007). Every action stays subject to validation: "Klem proposes, you validate."
  • A native MCP architecture (IA-006), which lets external AI agents operate on the client's carbon data in an auditable way.
  • A complete, tamper-evident audit trail (CERT-003), tracking every entry, change and validation with a timestamp and an identified author.
  • FAQ

    Can ChatGPT produce a reliable carbon footprint for a company?

    ChatGPT can help draft a methodology note or explain a concept, but it does not natively hold a sourced, versioned emission factor database, nor a structured memory of a company's activity data. Without a verified connection to reliable data and factors, a result produced by ChatGPT alone is not defensible in front of an auditor or a client.

    Can an LLM like ChatGPT or Claude be used to prepare a CSRD report?

    An LLM is genuinely useful to write, summarize or simplify CSRD content, as detailed in our CSRD report prompts guide. It does not replace the underlying footprint calculation (ESRS E1), which requires sourced data, a reproducible calculation under the GHG Protocol, and an audit trail a third party can verify.

    What is an AI hallucination applied to a carbon figure?

    A hallucination is an answer generated by an LLM that sounds plausible but is not grounded in a verifiable source. According to a 2025 study by OpenAI researchers, this phenomenon partly stems from training mechanisms that reward a confident answer over an admission of uncertainty. A hallucinated emission factor has the same numeric shape as a real one, without its source.

    Why must a carbon calculation be reproducible?

    A carbon footprint is compared year over year and reviewed by third parties (auditors, clients, regulators). If the same dataset and the same emission factor produce different results depending on when they are generated, comparison and verification become impossible. A deterministic calculation engine guarantees that the same input always produces the same output.

    What is the difference between an AI chatbot and an AI carbon agent?

    A chatbot answers questions in a conversation. An AI carbon agent acts on real data through tools: it can extract an invoice, match an accounting line to an emission category, or flag an inconsistency, while leaving final validation to a human. This distinction is detailed in our article on the AI carbon agent.

    Does connecting an LLM to a database solve these limitations?

    Partly. Connecting an LLM to a sourced factor database and structured activity data reduces the hallucination risk and improves traceability. A deterministic calculation engine, multi-entity boundary management, and an audit trail are still needed to turn a generated answer into a figure that is defensible in front of a third party.

    Conclusion

    An LLM alone writes and explains well, but it lacks structured activity data, sourced emission factors, a reproducible calculation, and an audit trail: the four elements a defensible carbon figure requires. The right architecture places AI on top of a reliable data system, with human validation before every action, rather than in place of that system.

    Before using AI for your carbon footprint, check that the answer relies on a traceable data source and a dated emission factor, not just a convincing sentence.

    Klem, Kabaun's AI assistant, acts on your carbon data under human validation → www.kabaun.com/en/klem

    Sources

  • Kalai, Nachum, Zhang (OpenAI), Vempala (Georgia Tech), "Why Language Models Hallucinate," arXiv:2509.04664, September 2025 · academic research. arxiv.org
  • GHG Protocol, "Corporate Standard" · international methodological reference. ghgprotocol.org