# What Kolibri-1's training disclosures establish, and what they do not

> Other vendors' models helped write, select and judge Kolibri-1's training data. The EU summary lists seven names; the technical report explains the different roles. These disclosures show how a model built for sovereign and specialised deployment can draw on other vendors' models during training. They do not establish which behaviours, if any, were inherited.

Source: https://counterproof.io/notes/kolibri-1-training-provenance/ · Published 5 October 2026 · CounterProof is a practice of Clavestra Capital Limited (Malta, C 113987).


**On 3 October 2026, Aleph Alpha published the weights of Kolibri-1, a German and English model.** It also released a technical report and the training-content summary that the EU AI Act asks of general-purpose model providers.

The report names the other companies' models that helped build Kolibri-1's training data and explains what each did. The summary lists seven of them. We read both to understand what the disclosures establish, what they leave open, and what the vendor means by "sovereign".

The important distinction is between roles. Some models supplied training text. Others helped decide which material was kept. Those roles describe different relationships with Kolibri-1.

## What the vendor disclosed

The EU summary asks: "If yes, specify the general-purpose AI model(s) used to generate the synthetic data if available on the market:" Aleph Alpha answers with seven names: Gemma-4-26B-A4B, Qwen3-32B, Qwen3.8-27B, Mistral-Nemo-12B, GLM-5.3, GLM-5.2-fp8 and Kimi-K2.6.

The technical report and model card describe what these models, and a few others, did. The table retains the documents' wording; its stage column follows the report's section structure.

| model | what the documents say it did | where in the training |
|---|---|---|
| Gemma-4-26B-A4B | rephrased seed documents, "using five prompt templates applied to each seed document" | pretraining data (report §2.3.2) |
| Mistral-NeMo-12B | "German synthetic rephrasings of web data were produced with Mistral-NeMo-12B" | stage not in the quoted card sentence (model card) |
| Qwen3-32B | scored "a fixed 1M-document English sample" of "our curated CC data" "using more than 20 prompts"; "Luxical-based distilled judge classifiers" were then used "when constructing our Kolibri dataset" | pretraining data (report pp. 28 to 29) |
| GLM-5.2, GLM-5.3, Qwen3.8-27B | "the main models we use to generate this data, and to regenerate parts of the open datasets" | post-training data, supervised fine-tuning (report p. 53) |
| GLM-5.2 | "writes two questions with six options" per document; a question enters training "only if GLM-5.2 solved it at least once and Kolibri at most six times" | long-context environment (report App. I.1.2) |
| Kimi K2.6 | "where Kolibri's attempts agree on a different answer or solve it at most once, Kimi K2.6 ... rules on the disputed options without seeing the key" | long-context environment (report App. I.1.2) |
| Kimi-K2.7 | "we apply Kimi-K2.7 ... to roll out every question in the environment, and these rollouts drive a filtering and repair pass" | the environment, before training (report p. 176) |
| Gemma-4-31B | "We keep a rollout when an LLM judge, Gemma-4-31B (Gemma Team 2026), accepts its answer, the gold answer appears in the retrieved text and the transcript ends on an answer" | retrieval environment (report p. 58) |
| Command A+ | "We keep a seed when its completions answer correctly on the first two contexts, abstain on Morgana's and pass a Command A+ ... judge" | retrieval and search data (report p. 58) |
| gpt-oss-120b | "we judge the whole conversation, including reasoning traces, with gpt-oss-120b ... to classify each as either carrying a CCP narrative, a refusal on general safety grounds, or being balanced. Conversations of the first kind are dropped." | post-training data curation (report p. 60) |

The report explains the purpose of the last row: "We acknowledge that several of our teacher models were built in China, so they carry known political biases on sensitive topics, while our model should adhere to the values of the democratic consensus in Europe."

It also describes three mechanisms that matter for any later attempt to recover training relationships from the model's outputs:

- **Filtering identity claims.** "A pattern filter drops every conversation in which the assistant claims to be another model or to come from another lab, unless it also names our own model, and ignores denials such as 'I am not Llama'".
- **Constraining scripts.** A reinforcement-learning format constraint states: "No script leak. No Chinese, Japanese and Korean (CJK) text leaks into a response that is otherwise non-CJK".
- **Varying reasoning openers.** In "reasoning-prefill distillation", a teacher's German reasoning begins with an opener such as "Gegeben ist:". The opener is "randomly picked per sample from a curated list per task domain, so that the model does not learn one fixed opening phrase".

## Names are not mechanisms

The EU template asks for one list under one heading. Its definition of synthetic data covers "model distillation or model alignment (e.g. AI feedback ...)", so Kolibri-1's seven names represent several kinds of relationship.

Rephrasing pretraining text and generating post-training data introduce another model's text into the training set. Scoring documents, judging rollouts and ruling on disputed items influence which material is kept. The documents do not establish whether these selection and judging roles leave traces in the trained model's outputs, or what those traces might be. We make no such claim.

This distinction determines what can be tested. Methods that try to recover a model's origins from its outputs look for inherited words, phrasing and habits. Relationships involving generated text are the ones such methods can hope to detect.

A disclosure that identifies both the mechanism and the training stage helps an independent researcher work out what is testable. The technical report provides that detail; the template would be more useful if it asked for it too.

## What "sovereign" covers

The report is titled "Kolibri: A Sovereign European Model on the Pareto Frontier". Its abstract explains what sovereign means in this context: "We built Kolibri for sovereign and specialised deployment, with a particular focus on German and on regulated domains such as public administration, industry, and aerospace."

The model card adds that the weights "are published by Aleph Alpha GmbH under Apache 2.0 license" and that "Aleph Alpha is a signatory of the EU GPAI Code of Practice".

These are statements about deployment and operation. The table describes how the training data was built. The report presents both.

Reading "sovereign" as a claim that no other vendor's model took part in training would misread the report. Treating those contributions as disproving the deployment claim would also go beyond what the documents establish.

## What the disclosure cannot establish

The disclosures identify candidate relationships. They do not establish which characteristics, if any, Kolibri-1 inherited from the models involved.

The report and summary also differ in coverage. Four models in the table (Gemma-4-31B, Command A+, Kimi-K2.7 and gpt-oss-120b) appear in the report but not in the summary's list. Two of the summary's seven names, Qwen3-32B and Kimi-K2.6, serve as a judge and an adjudicator rather than as generators of text.

The training figures come in two forms. The report's abstract gives 24 trillion tokens "across pre-training, mid-training, and long-context extension"; the model card breaks that down as "20T tokens of a filtered, bilingual corpus (~62.5% English, ~23.9% German, ~13.6% code)" and "Additionally trained on 3.44T in mid-training and 201B for long-context extension".

The predecessor, Kolibri Origin, is described as "an earlier iteration that validated our full pipeline at scale" and "is not released". A third party therefore cannot run it. Table 28 compares its post-training benchmark scores with Kolibri's.

The three mechanisms described above also limit the evidence available from outputs. Identity claims and stray scripts are excluded as evidence by those mechanisms; attempts to recover the relationships must work with subtler features.

These limits are part of what makes the case useful to us. Our research on model provenance holds that similarity between models is not descent. A method that claims to reconstruct relationships must be tested where the answer is known. Kolibri-1 offers a vendor-published list of candidates.

We are preparing a pre-registered test of whether the disclosed relationships can be recovered from outputs at all. We will publish the design, including its stopping rules, before running any model, and report the result in the design's own terms. This note is a reading of documents, not an experimental result.

## What a disclosure like this is for

Buyers in regulated sectors are starting to ask where a model came from. A useful answer should explain which models wrote its training text, which selected it, which judged it, and at what stage each was used. The EU template asks for a list; Kolibri-1's report shows that a fuller account can be given.

For researchers, that account supplies a known candidate list against which provenance methods can be tested. It offers firmer ground than inferring sources from behaviour and then arguing over the inference, where much of the field still stands.

We have no relationship with Aleph Alpha and have not spoken to them. This note draws on the public technical report, model card and EU summary. The quotations come from the locations listed below. We ran no model and made no measurement. We will correct any misreading brought to our attention.

## Sources

- Aleph Alpha, "Kolibri: A Sovereign European Model on the Pareto Frontier", technical report, 2026: https://aleph-alpha.com/downloads/tech-report.pdf. Locations: abstract (sovereign deployment; Kolibri Origin); §2.3.2 (Gemma-4-26B-A4B rephrasing); pp. 28 to 29 (Qwen3-32B scoring and the distilled classifiers); p. 53 (post-training generators); p. 54 (reasoning-prefill distillation); p. 58 (retrieval rollouts, Gemma-4-31B and Command A+ judges); p. 59 (teacher bias); p. 60 (identity pattern filter; gpt-oss-120b classification); the reinforcement-learning format constraints (script leak); Appendix I.1.2 (long-context environment: GLM-5.2 questions, Kimi K2.6 ruling); p. 176 (Kimi-K2.7 rollouts); Table 1 (Kolibri Origin not released).
- Aleph Alpha, Kolibri-1 model card: https://huggingface.co/Aleph-Alpha/Kolibri-1 (licence; language mix; Code of Practice; Mistral-NeMo-12B rephrasings).
- Aleph Alpha, "Public Summary of Training Content for General-Purpose AI models", Kolibri-1, the template under Article 53(1)(d) of Regulation (EU) 2024/1689: https://aleph-alpha.com/downloads/Kolibri_1_-_Sufficiently_Detailed_Summary.pdf (§2.5).
- Our research on provenance and reviewer independence: https://research.counterproof.io/agentic-stemmatics.md


## Questions this page answers

**Did Aleph Alpha train Kolibri-1 on other companies' models?**

Its report describes training its own weights; it does not describe using another model's weights. Other vendors' models helped build the training data: they rephrased web text, scored documents, generated and regenerated post-training data, judged rollouts and ruled on disputed items. The technical report describes these roles, and the EU training-content summary lists seven of the models involved.

**Does that contradict the 'sovereign' label?**

The report uses 'sovereign' for the deployment it is built for: open weights under Apache 2.0, with a particular focus on German and on regulated domains. That is a statement about operation. Which other models took part in building the training data is a separate statement, and the report makes both.

**Can you tell from a model's outputs which models took part in its training?**

That is an open question, and it can only be tested where the answer is known. Kolibri-1 is a case where the vendor published the candidate list. We are preparing a pre-registered test and will publish its design before any model is run.

**What should a buyer ask a vendor?**

Ask which models wrote the training text, which selected it, which judged it, and at what stage each was used. The EU template asks for a list of names. Kolibri-1's technical report explains the roles behind them.

