commercetools catalog migration

Description

Migrates a product catalog from any source system into commercetools through a canonical NDJSON feed contract — the pipeline never parses the source, so each engagement writes only a small adapter. Covers the feed contract, writing an adapter, ProductType derivation and its irreversible attribute constraints, decimal-to-minor-unit money conversion, project-wide slug allocation, category order hints, and the offline audit gate. Use when migrating or bulk-loading a product catalog into commercetools, writing an adapter for a source export, deciding SameForAll versus CombinationUnique, converting prices to centAmount, or diagnosing rejected Import API operations. Covers loading a catalog from a source system into commercetools — not the Classic-to-Modular catalog model migration of an existing project, which is the staged Modular Catalog migration guide. Not for ongoing sync after cutover, and not for customers, orders, carts, inventory, or promotions.

Installation

Recommended: install the full commercetools plugin. It includes this Skill, every other commercetools Skill, our pre-tuned Subagents, and the commercetools Knowledge MCP — which gives AI live access to the commercetools docs, GraphQL/OpenAPI schemas, and query validation. You only install once; every Skill on this site becomes available in every session.
Install the plugin

In any Claude Code session, run the following commands one at a time:

  1. Add the commercetools AI marketplace:

    /plugin marketplace add commercetools/commercetools-ai-plugins
    
  2. Install the commercetools AI plugin (Skills + Subagents + MCP):

    /plugin install commercetools@commercetools
    
Reload plugins

If you've updated the plugin or installed it in another window and need the current session to pick up the latest version:

/reload-plugins
Claude Desktop
Customize -> Personal plugins -> Create plugin -> Add marketplace -> Add commercetools/commercetools-ai-plugins. Then, click on the plugin and click Install.

Instructions Included

SKILL.md

commercetools catalog migration

Migration is one-time and decision-heavy. The deliverable is a reviewed set of decisions, not a loaded dataset: what was translated, what was approximated, and what was lost, with a rationale per decision. A migration claiming zero information loss is not being honest.

If the work is a recurring sync from a PIM, ERP or OMS, this is the wrong approach: that is integration, contract-heavy rather than decision-heavy.

The source boundary is a data contract, not a code module

any source  →  adapter  →  canonical catalog feed  →  pipeline  →  commercetools
               per          NDJSON, schema-checked    reusable
               engagement   THE CONTRACT
The pipeline — commercetools/ct-catalog-migration-pipeline, cloned beside the engagement — never parses a source system. It defines a feed contract and starts there. Everything source-specific lives in the adapter, upstream of the contract.
This is not decoration. A universal parser for any source platform does not exist — no two installations agree on their type system — so source knowledge inside the pipeline is what destroys retargetability. A seam made of data rather than code makes that coupling impossible rather than discouraged. The contract is schema/catalog-feed.schema.json in that repository; treat it as the API.

Before you start

  1. Docs search (required, run first) — The first time you use this skill in a session you must run this before answering. It gathers the latest verified documentation as your primary grounding, filtered to the products this skill covers. Use this script for documentation search while working with this skill; the Knowledge MCP covers everything else. Always confirm details against retrieved documentation rather than the skill text alone:
    node scripts/docs-search.mjs \
      --query "<extract key terms: Import API, ProductType attributes, catalog model, Standalone Prices, Category slug>" \
      --app-name "<host app: claude-code, claude-chat, cursor, codex, copilot — or the host's own name>" \
      --model "<current-model>" \
      --limit 10
    
    The limits and constraints in Critical were verified in September 2026. Re-check any of them here before relying on one.

Workflow

  1. Interview, then write the config. Do not hand over a template to fill in. Ask the questions, then write migration.config.json from the answers — the config is the record of a conversation, not a form.
    Every question, what it writes, and why none of them can be defaulted: references/step-0-interview.md. Work through it in order — the order is load-bearing:
    1. Size the export (minutes): file list, line counts, largest variant count. Structural only — no mapping, no types, no decisions.
    2. Ask 0a, the catalog-model question, with that number in hand. Both its answers can end the engagement, and both are cheaper to learn now than after an adapter exists. A catalog over 100 variants per product cannot be loaded under Classic at all.
    3. Discover the source (0b) — references/source-discovery.md — then ask the rest (0c, 0d), because the data answers some of them.
    Then open the decision log. The config records what was chosen, not who chose it, what the alternative would have cost, or that a question was asked at all. Write one entry per decision from 0a and 0c, plus any question that could not be answered — an open question with an owner is a plan, an unasked one is a surprise. Format and placement: references/decision-log.md.
    Keep writing it at every step from here, not at the end: a log reconstructed afterwards records what someone remembers deciding, reliably the subset that turned out well, while the entries that matter are the irreversible ones taken under uncertainty.
  2. Write the adapter — one product first. Read the source, emit NDJSON against the contract. This is the only code an engagement normally writes, and it is deliberately disposable. Work through references/writing-an-adapter.md against a real export — one designed against an assumed source model gets rewritten — with the record types in references/catalog-feed-contract.md.
    Take the thin vertical slice before writing the rest: one product, two or three variants, through validate, derive, plan and audit. Writing every SKU in one pass usually gets away with it, which is why this step keeps getting skipped — and is not a reason to skip it. The axis model, the money conversion and the attribute constraints are wrong in the same way across five records as across five thousand, and they are irreversible once loaded.
    Log every point where the adapter decides rather than translates: a code derived from a display name, a hierarchy level collapsed, an inherited price resolved, a locale tag normalised, a sentinel skipped or interpreted, a field dropped, and every silent cap — anything sampled, truncated or skipped. These are the decisions most often lost, because they are made while writing code and feel like implementation details. They are not: they are what the migration did to the customer's data.
  3. Run the offline stages in order. Each gates the next, so a defect is reported at the earliest stage that can see it:
    cd ct-catalog-migration-pipeline
    npm run pipeline -- validate --config ../migration/migration.config.json
    npm run pipeline -- derive   --config ../migration/migration.config.json
    npm run pipeline -- plan     --config ../migration/migration.config.json --payloads
    npm run pipeline -- audit    --config ../migration/migration.config.json
    
    None of these need credentials, by design: a broken feed or an unloadable plan reports the real problem on a laptop with no .env.
    Log the findings you accept rather than fix. A diagnostic that changes the adapter leaves no decision behind — the feed is regenerated and the finding disappears. One accepted is a decision: 14 disambiguated slugs and the URLs that result, a products-in-no-category warning that is correct for this catalog, an approaching-variant-limit acknowledged with a plan to revisit. Record the count and the artefact, not every finding.
  4. Read out/MODEL-REVIEW.md before loading. It is the sign-off artefact: irreversible choices, information loss, and anything guessed.
    Log the review outcome, not its contents — one entry naming what was accepted and against which plan. MODEL-REVIEW.md is regenerated on every run and can hold thousands of entries; copying them into an append-only log guarantees the two disagree by the next run.
  5. Check the target project once credentials exist. preflight answers whether the project will accept the plan, in a handful of GETs, before anything is written:
    npm run pipeline -- preflight --config ../migration/migration.config.json
    
    --apply writes to the customer's project, and so does any manual change made to get past preflight. Log every one: project key, timestamp, exactly what changed, who authorised it. --apply reports project-updated to stdout and nowhere else, so an unlogged change is one nobody can account for. Manual changes count doubly — adding a country to make country-scoped prices load affects shipping and tax, which is why preflight refuses to do it.
  6. Load, dry run first. load sends nothing without --execute, and writes the exact request bodies to out/load-requests.json for review:
    npm run pipeline -- load --config ../migration/migration.config.json
    npm run pipeline -- load --config ../migration/migration.config.json --execute
    
    Log one entry per --execute: project key, timestamp, what was sent, the outcome. out/load-result.json holds the detail but is regenerated and gitignored.
  7. Verify, and do not treat the load's own report as the answer. An accepted Import Request is not an imported resource — the API validates asynchronously, so a load can report every request accepted while individual operations are rejected afterwards. verify reads the project back and reconciles it against the plan. Every call is a GET:
    npm run pipeline -- verify --config ../migration/migration.config.json
    
    It needs view_products and, under standalone pricing, view_standalone_prices — reading is a different scope set from writing.
    Log the result as the closing entry for the run, including a clean one. "Reconciled, no differences" is the sentence someone needs months later when asking whether the load was ever checked at all.
  8. Fix defects in the adapter, never in the feed by hand. The feed is regenerated on every run.

Critical

The failure modes that are silent or unrecoverable, each stated in full in the reference beside it — this list exists so none is missed. Verified against public documentation in September 2026; re-check any limit before relying on it.
Irreversible, so decide before the first load
  • Attribute constraints are a one-way door. changeAttributeConstraint accepts only None: a constraint can be relaxed, never tightened or switched. A wrong SameForAll / CombinationUnique means recreating the attribute and rewriting its data. → product-model.md
  • A product selection's mode is fixed at creation — there is no changeMode action, and re-importing the other mode reports imported and changes nothing. → catalog-feed-contract.md
  • A Product's ProductType cannot be changed after creation. A project loaded before ProductType keys were prefixed trips product-type-keys-unprefixed and cannot be moved onto the new keys by re-running. → running-the-pipeline.md
  • productCatalogModel is fixed at map time. Classic embeds variants (1–100) and allows embedded prices. Modular makes them standalone VariantImport resources (to 10,000), allows only standalone prices, and cannot import defaultVariant at all. Changing it means re-running plan. → product-model.md
Silent when wrong — no error, wrong catalog
  • Money is minor units, and the digit count is not 2 everywhere. JPY and KRW are 0; BHD, KWD, OMR and TND are 3. Defaulting to 2 multiplies prices by 100 without erroring. Convert on the digit string, never through a float. → pricing-and-money.md
  • A localized value must never be a variant axis. Identity would key differently per language and break the moment a second locale is added. axisValues takes codes; axisLabels takes display text. → catalog-feed-contract.md
  • isSearchable must agree across every ProductType sharing an attribute name, or that attribute leaves search, filters and facets everywhere — and the import still succeeds. → product-model.md
  • Omitted fields are deleted on import. ProductVariantImport and StandalonePriceImport drop whatever you leave out, so an update must resend everything worth keeping. ProductVariantPatch is the only partial update. → running-the-pipeline.md
  • An unresolved KeyReference expires after 48 hours, taking its price with it: a green load and a partly priced catalog. → running-the-pipeline.md
  • publish: false unpublishes. It is an instruction, not "leave publication alone", and the pipeline plans it on every product. Correct on a first load; on a re-load into a live project it takes every product it touches offline, with every operation still reporting imported. → running-the-pipeline.md
Project-wide, not per resource
  • An attribute name holds exactly one type per Project, not per ProductType, so a self-consistent plan can still be refused by what the project already holds — AttributeDefinitionTypeConflict, then AttributeNameDoesNotExist on every product using it. No update action changes a type: the fix removes and re-adds the attribute, deleting its values on every product that has one. Enum values may differ freely. preflight is the only stage that can see it. → product-model.md
  • Slugs are unique per locale across the Project. Retail taxonomies reuse names, so collisions are normal — and resolving one changes a URL, which is information loss and somebody else's decision. → categories-and-slugs.md
Shape of the model
  • Prices are scoped by channel; a store only chooses which channels apply. There is no store field on a price. Two distribution channels on one store, both priced, is legal and ambiguous — the storefront has to call setLineItemDistributionChannel. → pricing-and-money.md
  • A store with no product selections offers everything; one whose selections are all inactive, with an Individual among them, offers nothing. → catalog-feed-contract.md
  • Channels and customer groups cannot be imported — load creates them through the platform API before any import, with keys used verbatim: the one exception to the prefix rule, and the one thing a prefix-scoped teardown leaves behind. A price channel needs ProductDistribution or the API refuses the price. productSelection is importable and prefixed; store is not, and runs last. → catalog-feed-contract.md
  • Import Resources accept only KeyReferences, never ids. Key everything <prefix>-<sourceCode>, ProductTypes included — though a ProductType's name comes from the unprefixed code, or the Merchant Center reads "Mig Apparel Basic". → catalog-feed-contract.md
  • An order hint is a string strictly between 0 and 1 and must not end in 0 — a zero-padded scheme emits 0.10 and is rejected. → categories-and-slugs.md
Method
  • Do not serialise the Import API on polling. Operations are retained 48 hours precisely so an unresolved reference resolves when its target arrives; polling each stage to completion serialises a pipelined design, and frequent summary polling slows the import. Retry only rejected — other states are retried internally.
  • Verify on a separate code path. A verifier sharing logic with the writer confirms the writer's bugs. The audit gate reads the written plan back from disk and re-derives every invariant from the drafts as data.
  • Refuse to write data you cannot express correctly. Skipping is recoverable; wrong data in production is not. Report it as loss instead.

Reference index

TopicReference
Profiling a source sample and producing a mapping proposal (step 0b)references/source-discovery.md
The engagement decision log: what to record at each step, and what not toreferences/decision-log.md
The feed contract: every record type, with the two load-bearing design rulesreferences/catalog-feed-contract.md
Writing an adapter for an arbitrary source, and what to establish firstreferences/writing-an-adapter.md
ProductTypes: declared vs inferred, constraints, levels, searchabilityreferences/product-model.md
Money, price scopes, validity windows, embedded vs standalonereferences/pricing-and-money.md
Categories: tree, slug allocation, order hintsreferences/categories-and-slugs.md
Running the pipeline: each stage, each gate, and what is still absentreferences/running-the-pipeline.md

Scope

In scope: ProductTypes, Categories, Products, Variants, and prices — the irreducible core of a loadable, browsable catalog.

Out of scope, and deliberately so. Ongoing sync after cutover and delta feeds as a running service are integration, not migration. Customers, orders, carts and payments are a different migration with different invariants. Promotions and discounts are usually a re-modelling exercise rather than a translation, and attribute modelling in the abstract, tax modes and pricing strategy are product data modelling, upstream of this. Inventory and supply channels, and cutover sequencing, freeze windows and rollback, are planned and not built.

Never migrate credentials, payment tokens or PSP references. They are not portable, and re-keying them silently breaks reconciliation with the provider. Passwords do not migrate either — commercetools cannot import a foreign hash.

Checklist

Before running the pipeline:

  • The catalog-model question was asked before any adapter work, and the largest product's variant count is known. A catalog over 100 variants per product cannot be loaded by this pipeline at all.
  • The config was written from answers, not copied and guessed at. Nothing in it still reads REPLACE-ME.
  • The source has been inspected, not assumed — the adapter is written against a real export.
  • Every axis value is a language-independent code, with display text in axisLabels.
  • market.currencyFractionDigits has an explicit entry per currency.
  • target.priceMode matches how the implementation prices. Standalone pricing also needs manage_standalone_prices on the API Client, which manage_products does not grant.
  • keys.prefix names this engagement, so teardown can scope itself. Every resource carries it, ProductTypes included. Channels and customer groups are the only exception: load creates the missing ones with their keys verbatim, and teardown will not remove them.
  • preflight reports no product-type-keys-unprefixed. If it does, the project was loaded before ProductType keys were prefixed, and a Product's ProductType cannot be changed after creation.
  • If the feed declares any customerGroup, the API Client has view_customer_groups and manage_customer_groups — manage_products grants neither, and load needs both to create one. Channels are covered by manage_products.
  • If the feed declares any productSelection, the API Client has view_product_selections and manage_product_selections; if it declares any store, view_stores and manage_stores. manage_products grants none of these.
  • Every productSelection mode has been confirmed with whoever owns the assortment. It is permanent, and the two modes are opposites.
  • No store's selections are all inactive — that offers no products, where an empty list would have offered every product.
  • DECISIONS.md exists, is committed, and has an entry for every step 0 decision — including the ones still open.

Before loading:

  • validate, derive, plan and audit all pass — in that order, in one sitting. If plan fails it writes nothing and the previous plan.json survives; audit refuses a plan whose feed digest no longer matches (plan-stale), but a plan with no digest at all can only report unknown, so do not rely on the check to catch a replay.
  • out/MODEL-REVIEW.md has been read and its irreversible choices accepted.
  • SameForAll and CombinationUnique are right — they cannot be tightened later.
  • Each recorded information loss is accepted, or the adapter is fixed.
  • Sample payloads have been eyeballed, money values checked against the source.
  • Slug changes forced by project-wide uniqueness have been shown to whoever owns SEO.
  • Attribute names shared across ProductTypes agree on isSearchable.
  • preflight reports no attribute-type-conflict. An attribute name may hold one type per Project, and the conflict can come from a ProductType the migration never touches.

After loading:

  • Counts and structural properties verified through a separate code path.
  • Products published deliberately, as a separate step — the pipeline always plans publish: false. If the target project already holds published products, confirm before --execute that unpublishing the ones this run touches is intended: publish: false unpublishes, silently.
  • The container summary was read inside 48 hours of the run. Past that, failed operations are deleted and a summary of survivors reads clean. verify reads resources, which do not expire.

References

catalog-feed-contract.md

The canonical catalog feed

The pipeline's source boundary. Adapters write these shapes; the pipeline reads nothing else. No source system is named anywhere in the contract, by design.

Authoritative schema: schema/catalog-feed.schema.json in the pipeline repository (JSON Schema draft 2020-12). This page explains the intent; the schema is what validates.

Shape

A feed is a directory of *.ndjson files. One JSON record per line, each tagged with _type. Splitting across files is the adapter's choice — record order is not significant, and forward references are allowed, so an adapter can stream products before the categories they belong to.
{"_type":"category","code":"tops","name":{"en-GB":"Tops"},"parent":"apparel","sourceOrder":1}
{"_type":"attributeDefinition","name":"colour","type":"lenum","level":"variant","axis":true,
 "values":[{"key":"BLK","label":{"en-GB":"Black"}}]}
{"_type":"product","code":"TEE-01","name":{"en-GB":"Cotton Tee"},"categories":["tops"],
 "axes":["colour","size"],"attributes":{"material":{"en-GB":"Cotton"}}}
{"_type":"variant","sku":"TEE-01-BLK-M","product":"TEE-01",
 "axisValues":{"colour":"BLK","size":"M"},
 "prices":[{"currency":"GBP","amount":"19.99","country":"GB"}]}

Variant axes

An axis is a dimension along which a product's variants differ. A tee in 3 colours × 4 sizes has two axes — colour and size — and each of its 12 variants is one point in that grid.
The term is the feed's own. commercetools has no axis field: the platform says the same thing with attributeConstraint: 'CombinationUnique' on a variant-level attribute, and derive performs that translation. The feed records what the source means; the pipeline chooses the mechanism.

Three records share the work:

  • an attributeDefinition with "axis": true declares that an attribute may carry variant identity — only legal at "level": "variant";
  • a product lists the axes it varies by in axes, in order; absent or empty means a single-variant product;
  • a variant gives its own coordinate in axisValues, with display text kept separately in axisLabels.

An axis is not merely an attribute whose value happens to differ between variants. Four things follow from marking one:

  • The constraint is a one-way door. CombinationUnique can be relaxed to None later but never re-tightened or switched, because changeAttributeConstraint accepts only None. Getting it wrong means recreating the attribute and rewriting its data, which is why derive records every such choice as irreversible.
  • Values must be language-independent codes, never localized text — the next section is that rule.
  • Coverage and uniqueness are enforced. Every variant must carry a value for every axis its product declares, and no two variants of one product may share a combination.
  • An attribute cannot be an axis in one product and a plain attribute in another product of the same ProductType. One definition carries one constraint, so the ProductType has to be split — see product-model.md.

Two rules that are load-bearing

These are enforced by the schema rather than left to discipline, because both failure modes are silent.

Axis codes and axis labels are separate fields

axisValues takes a language-independent code per axis. axisLabels takes display text. A localized value therefore cannot become variant identity.
// Correct — identity is a code, display text is separate
{"axisValues":{"colour":"BLK"},"axisLabels":{"colour":{"en-GB":"Black","de-DE":"Schwarz"}}}

// Wrong, and the schema rejects an ltext axis: the same SKU would key
// differently per language and break when a second locale is added
{"axisValues":{"colour":"Black"}}
If the source only has a localized colour name with no stable code, derive one in the adapter and record the derivation. That derivation is a decision worth reviewing, not an implementation detail.
One code, one label. A lenum holds a single label per key, so if two products give the same axis code different display text — M as "Medium" on one style and "Mid" on another — only the first survives. derive reports which labels it dropped (lossy, needing review) rather than picking quietly, because the winner used to depend on feed line order.
Three ways out, and the choice turns on whether the axis has to work as a cross-product facet:
What it costs
Pick one label deliberatelyThe other wording is lost. Fine when the labels are synonyms ("Medium" / "Mid"); wrong when they name different things.
Make the codes differKeeps both meanings and keeps the facet — but only if the source has something stable to differ by. Deriving a code from a field the source does not guarantee is stable breaks a repeatable load, so check the re-export answer from step 0c first.
Keep the shared code, carry the exact wording as its own attributeKeeps the facet and every label. The per-variant text stops being variant identity, which is usually correct — it was display text, not a code.
What to avoid: keying the axis on something per-product, such as a full style-plus-colour code. It resolves the collision and destroys the axis as a facet — every product gets its own value set, so "show me everything in black" cannot be expressed. It is the route a source with no distinguishing code invites, and the third row above is the better answer in exactly that situation.
axisLabels only reach the project on an enum or lenum axis. Those are the only types with a key and display text as separate fields. An axis declared text is its own display text, so every label supplied for it is discarded — derive warns (axis-labels-ignored) rather than dropping them silently.

Money is a decimal string, never a JSON number

{"currency":"GBP","amount":"19.99"}   // correct
{"currency":"GBP","amount":19.99}     // rejected by the schema
A binary float cannot represent most decimal amounts exactly, and the error surfaces as an off-by-one cent on a fraction of records — the worst kind of migration defect, because it looks like everything worked. Minor units are derived at map time from the currency's configured fractionDigits, and an amount with more decimal places than the currency allows is a hard error rather than a silent rounding. Detail in pricing-and-money.md.

Record types

All code fields are language-independent identifiers matching [A-Za-z0-9_-], and they end up inside commercetools keys, so they must be stable across exports. A code that changes between runs creates a duplicate rather than an update.

channel and customerGroup — prerequisites, not imports

These two are unlike every other record: they are created through the platform API, not the Import API. The Import API has no channel or customer-group resource — neither can be imported at all.
That matters because prices reference them. A price scoped to a channel the project does not hold becomes an Import Operation that sits unresolved for 48 hours and then expires, taking the price with it. The load reports every request accepted, the operation is gone before anyone looks, and the price is simply absent.
So load creates them first, before any import, and only the ones that are missing:
  • it reads the project for the declared keys;
  • absent ones are created through the platform API;
  • present ones are left exactly as they are — never patched;
  • a present channel lacking a role the plan needs is an error, reported and not modified, because roles govern stores and inventory too.
They are the first two entries in loadOrder, and they are reported separately from the containers because they are a different mechanism: created synchronously, with no container and no operation states. The dry run reads the project as well, so it can name exactly which ones it would create; if that read fails it says so and reports them as existence unknown rather than guessing.

Three consequences follow:

  • The code is the key, used verbatim — never prefixed. A price references the project's own channel key, and channels are often created by store setup long before a catalog migration runs. These two are the only deliberate exception to <prefix>-<sourceCode>. ProductTypes used to be a second, accidental one — keyed verbatim from the feed's productType code — and are now prefixed like everything else.
  • A teardown scoped to keys.prefix will not remove them. If the load created one, it has to be removed by hand.
  • Creating them needs scopes the import does not. manage_products covers channels; customer groups need view_customer_groups and manage_customer_groups, which manage_products does not grant.
channel fieldNotes
coderequired — the channel's key in the project, verbatim
rolesrequired, at least one. A price-scoped channel needs ProductDistribution
name, descriptionoptional, localized
customerGroup fieldNotes
coderequired — the key in the project, verbatim
nameoptional; CustomerGroup.name is required by the API, so the code stands in
ProductDistribution is not cosmetic. The API refuses a Standalone Price referencing a channel without it — MissingRoleOnChannelError, "does not have the required role" — so validate makes that an error under priceMode: standalone and a warning under embedded, where the price imports but channel price selection cannot find it.
validate refuses a price scoped to an undeclared channel or customer group, and warns about a declared channel nothing references. preflight then reads the project: a missing one is a warning — load will create it — while one that exists with insufficient roles is an error, because that is the case load refuses to fix for you. preflight --apply creates neither; it changes project settings only, and creating resources belongs to load.

productSelection and store — assortments, and where they apply

These two answer a different question from everything else in the contract: not what the catalog contains, but which products exist, and at what price, in a given shopping context. They also split along the line that organises this whole pipeline:
Import APIKeyTeardown
productSelectionyes, resource product-selection<prefix>-<code>removed by a prefix-scoped teardown
storeno — platform APIverbatimleft behind

How prices, channels and stores actually relate

There is no store field on a price. Ever. The relationship is indirect and mediated entirely by channels:
Price ──carries──> Channel (role: ProductDistribution)
                      ^
Store ──lists─────────┘  distributionChannels[]
In a store's context the candidate prices are those whose channel the store lists, plus every price with no channel at all. Selection among the candidates then follows the documented fallback: customer group + channel + country, then progressively dropping dimensions, down to the bare price.

Three consequences worth planning around:

  • The channel is the pricing unit; the store is only the binding. Giving three stores three different prices means three channels. A store with no distribution channels sees only channel-less prices.
  • Equal specificity is ambiguous, and the platform will not resolve it. If a store lists two distribution channels and a variant is priced in both, both match at the same fallback step, and the storefront has to disambiguate with setLineItemDistributionChannel. Modelling "one channel per price list and one per store" is legal and needs middleware — know that before choosing it.
  • Cheapest wins within a step. A stray channel price silently undercuts the intended one rather than erroring.
supplyChannels is the inventory twin — the same Channel resource with the InventorySupply role. It is carried for fidelity and is the one field this pipeline cannot follow through on: no inventory is imported, so the channels are created and wired and every one of them holds zero stock. validate warns (store-inventory-not-migrated) rather than letting a correct-looking store imply migrated stock.

Product selections decide which products exist

A selection is a named subset of the catalog. On its own it does nothing — it takes effect only when a store references it.
productSelection fieldNotes
coderequired
namerequired, localized
modeIndividual (allowlist) or IndividualExclusion (denylist). Defaults to Individual. Permanent
The mode is fixed at creation, and re-importing does not say so. The update actions are setKey, changeName, add/exclude/remove product and the variant setters — there is no changeMode. Importing a different mode onto an existing key reports imported and leaves the mode untouched: nothing fails, nothing warns, and the assortment means the opposite of the plan. The API does not document this, so treat it as observed behaviour rather than a guarantee — but do not design around the import fixing a mode. It belongs in the same class as attributeConstraint: preflight reports a conflict with an existing selection as an error, verify reports drift as "delete and recreate", and plan records the choice as irreversible. Choose by which list is shorter to maintain, not which is shorter today.
Membership is authored on the product, not on the selection:
{"_type":"product","code":"TEE","selections":[{"code":"uk-assortment"}]}
{"_type":"product","code":"DRESS","selections":[{"code":"uk-assortment","includeSkus":["D-S","D-M"]}]}
selections[] fieldNotes
coderequired — a declared productSelection
includeSkusonly these variants
excludeSkusall variants except these. Mutually exclusive with includeSkus
With neither, every variant is in. A SKU that is not a variant of that product selects nothing and the platform prunes it silently, so validate makes it an error.

Two things follow from the Import API's shape, and they are the reason the feed is authored this way round:

  • Assignments cannot be split. The API replaces omitted fields, so all of a selection's assignments travel in one resource. plan inverts the product-side authoring and assembles the complete list per selection; there is no batching to fall back on, and an assortment of twenty thousand products is one large request. plan flags anything above a thousand for review.
  • One line per record stays true. A selection carrying its assignments in the feed would be a single twenty-thousand-element line.

Stores decide where it all applies

store fieldNotes
coderequired — the store's key in the project, verbatim
nameoptional, localized
languages, countriesoptional; must be a subset of the project's, checked by preflight
distributionChannelscodes of declared channels, each needing ProductDistribution
supplyChannelscodes of declared channels, each needing InventorySupply
productSelectionsat most 100, each {code, active}; active defaults to true
Deliberately unsupported: storefront URLs and custom. Neither is catalog data, and neither can be validated or verified here.
The activation rule reads backwards, so read it twice:
Store's productSelectionsProducts offered
emptyall of them
at least one activeonly the active ones' contents
all inactive, ≥1 Individualnone at all
all inactive, only IndividualExclusionall of them
An empty list is permissive; a list that is switched off is not. validate makes the third row an error (store-exposes-no-products), because it is the one an author reaches by accident while trying to stage a rollout.

Loading, and why the order inverts

A store references selections that the Import API creates asynchronously, and a store cannot be created pointing at one that does not exist yet. So store is a platform stage that runs last, after every import — not alongside channel and customer-group, which must run first.
load refuses without --wait when a store references selections. Without waiting there is no moment at which the store stage is safe, and refusing up front is cheaper than a catalog whose stores were never wired.
An existing store is never modified. setDistributionChannels and setProductSelections replace the whole array, so applying a plan over a store the project already configured would discard wiring this migration knows nothing about. preflight and load both report the difference and stop.

category

FieldNotes
coderequired
namerequired, localized
parentcode of the parent; omit for a root. Forward references allowed.
slugoptional; derived from name when absent, then disambiguated
sourceOrderinteger sibling order — prefer this over orderHint
orderHintonly if the source genuinely has a valid (0,1) string
descriptionoptional
externalIdoptional — the only externalId that reaches commercetools; CategoryImport has the field
Supply sourceOrder and let the pipeline encode a legal order hint. See categories-and-slugs.md.

attributeDefinition

Emit these when the source declares a type system. When the feed carries none, the pipeline infers definitions from observed values and forces an explicit review — inference is a fallback, not the happy path.

FieldNotes
namerequired, language-independent
typetext, ltext, enum, lenum, number, boolean, date, datetime, time, money, reference
levelproduct (invariant across variants) or variant
axistrue when part of variant identity; only valid at variant level
settrue to wrap the type in a set
valuesrequired for enum/lenum
searchableoptional; overrides productTypes.searchableByDefault for this attribute
required, label, unitoptional
searchable is per attribute, and the config value is only the fallback. A source that declares searchability itself — items.xml's search="true|false", for instance — should pass it through rather than let every attribute inherit one project-wide default. Fixture: declared-searchable.
Whatever is set has to agree across every ProductType that shares the attribute name, or the attribute becomes unavailable for search, filters and facets everywhere — see product-model.md.
unit has no commercetools equivalent and is carried into the label, which loses machine-readability — recorded as information loss. It applies to any type, not only number.

product

FieldNotes
code, namerequired
productTypegroups products sharing an attribute set; omit to take the configured default
axesordered attribute names distinguishing this product's variants; absent or empty means a single-variant product
attributesproduct-level values, invariant across variants
categoriescategory codes
slug, descriptionoptional
externalIdoptional, feed-only — dropped at map time (see below)

variant

FieldNotes
skurequired, unique across the whole feed
productowning product code; forward references allowed
axisValuescode per axis; every axis the product declares must be present
axisLabelsdisplay text per axis; never used for identity
attributesvariant-level values
pricesevery price explicit — see below
imagesurl required — relative or absolute, as the source stores it. width/height are optional but default to 0x0, recorded as information loss: a storefront that reserves layout space from the declared size cannot, so supply them whenever the source has them
assetsmedia assets — prefer these over images when the source has several renditions per shot
isMasteroptional hint; at most one per product. Absent ⇒ lowest SKU by sort order — deterministic, but arbitrary as merchandising. See below
keyoptional
externalIdoptional, feed-only — dropped at map time (see below)

Relative image URLs

image.url accepts a relative path. Emit what the source holds rather than resolving it: the host is usually not in the export, and the contract used to require an absolute URI, which forced adapter authors to invent one.
A feed with any relative URL needs media.baseUrl in the config. validate refuses the pair without it — naming the count, an example and the decision — and plan resolves against it via RFC 3986, so a root-relative /medias/x.jpg ignores any path on the base. Absolute URLs are untouched, and mixing is fine.
Assets follow the same rule — see asset below. Both are media, and a second media path with its own resolution is how one of them ends up wrong.

asset

On a variant or a category. An asset holds several sources, which is the reason to prefer it over images: an image carries exactly one URL, so a source with thumbnail, product and zoom renditions has to discard two of them. A hybris MediaContainer, or any format table, is one asset with one source per format.
FieldNotes
coderequired, stable — becomes <prefix>-<code>, because Asset.key is required by the API. Unique per variant or per category, not per project — see below
sourcesrequired, at least one
nameoptional, localized — derived from code when absent, and recorded as information loss
descriptionoptional, localized
tagsoptional strings

Each source:

FieldNotes
urirequired; relative or absolute, resolved exactly as an image URL is
keyoptional — what tells renditions apart, e.g. zoom. Unique within the asset
contentTypeoptional
width, heightoptional; both or neither, emitted as dimensions. Absent means 0x0, with the same cost as an image's — see the images row
The uniqueness scope is the owner, not the project. The API is explicit: Asset.key "is unique per Category or ProductVariant". So one source asset reused across the sizes of a colour is one code on several variants, which is legal and usually what you want — the source's own container code can go straight through. Only two assets on the same variant may not share a code (duplicate-asset-key).

This is worth stating because the opposite assumption — asset keys unique project-wide — makes the natural mapping look illegal and pushes an author into inventing per-SKU codes to satisfy a rule that does not exist. An over-strict reading here does not merely add noise; it changes the shape of the data.

The audit gate checks the asset key like any other key — charset, and no two assets on one owner claiming one — plus at least one source, a name in some locale, and no duplicate source keys inside an asset. verify compares assets by key and source URI, because an asset pointing at the wrong file renders a broken image and nothing else reports it.

Resolve in the adapter what commercetools cannot express

Price inheritance. commercetools does not inherit prices between variants. If the source inherits down a hierarchy, resolve it in the adapter so every SKU carries an explicit price. Most target prices usually exist because inheritance was resolved rather than skipped.
Variant hierarchies deeper than one level. commercetools has exactly one variant level. Collapse an N-level source hierarchy in the adapter and emit a flat SKU list with explicit axis values. Because the contract carries variants flat, the pipeline's own types never encode a hierarchy shape — which is what lets a 2-level and a 4-level source share the same pipeline.
Anything that needs a different commercetools concept. A quantity break is a Cart Discount; a net/gross decision is a tax mode; a category that is really a facet is an attribute. This pipeline's job is to detect and report that the source needs one, not to invent it.

Choosing a code

The code becomes part of a key, so:

  • Prefer a stable machine identifier over anything a merchandiser can edit.
  • Never use a localized or display value.
  • Preserve the raw source identifier in externalId as well — but read the next section first, because only a category externalId survives.

externalId only maps on a category

commercetools has externalId on Category and nowhere else in this pipeline's reach: CategoryImport has the field, ProductDraftImport and VariantImport do not. So a product or variant externalId is accepted by the schema, carried through the feed, and dropped at map time.
validate reports it as external-id-not-mapped rather than letting it go quiet. The field is still worth setting as feed provenance — it records what the source called this record — but it will not reach the project.
If a downstream system has to join on it, declare it as an attribute and emit it as one. An ERP article number or PIM id belongs in attributes: { articleNumber: "..." } with an attributeDefinition to match; that is the only form that lands.
This one is worth double-checking rather than assuming: setting externalId on a product or variant looks like it works at every stage — the schema accepts it, the feed carries it, and nothing warns until validate reports external-id-not-mapped. It survives casual review for exactly that reason.

Validation

validate runs in two phases. Schema conformance first; catalog-wide integrity only once every record parses, because a rejected record drops a product from the index and makes the integrity pass invent orphans that are really just consequences.

Integrity covers: duplicate codes and SKUs, missing category parents, category cycles, orphan variants, products without variants, missing category references, axis coverage and uniqueness, localized axes, attributes written but not declared, level mismatches, and multiple master-variant claims.

Every diagnostic names a file and a line. Fix the adapter, not the feed.

categories-and-slugs.md

Categories, slugs and order hints

Reference: Categories.
Limit: 10,000 categories per project (soft; increasable after a performance review — see limit increase guidance).

The tree

Categories carry a single parent. The feed allows forward references, so an adapter can emit in any order; the mapper emits parents before children so the import can also be streamed in that order.
Cycles and missing parents are caught by validate; a parent reference that survives into the plan but names nothing is caught by audit. An unresolved KeyReference would otherwise sit in the Import API for 48 hours and then expire.

Slugs are unique across the Project per locale

The same slug may repeat across locales of one category, but not across two categories. Pattern: [A-Za-z0-9_-]{2,256}. Indexes cover the first 15 project languages, so slug queries beyond that are slower.
Retail taxonomies reuse names freely — "Accessories" under both Menswear and Womenswear — so collisions are normal and have to be resolved rather than reported and abandoned.

Resolution appends the source code, which keeps the result traceable to the record that produced it:

mens-acc    →  accessories
womens-acc  →  accessories-womens-acc

This changes the URL the migration produces, so it is recorded as information loss. Generate the slug report from the plan, not from the loaded project: the old URL structure only exists before cutover, and the SEO owner needs to see the changes forced by project-wide uniqueness while there is still time to argue.

Derivation from a name strips diacritics by decomposing them rather than dropping the letters, so Oberbekleidung für Männer becomes oberbekleidung-fur-manner rather than a run of dashes. A name that slugifies to nothing — punctuation only, or a non-Latin script with no transliteration — falls back to the source code.

Product slugs follow the same uniqueness rule and the same pattern.

Order hints

An order hint is a string holding a decimal strictly between 0 and 1, and it must not end in 0.
Both halves matter, and the second one catches people out: the natural implementation is to zero-pad an index so the values sort correctly, but that produces 0.10 for the tenth sibling, which is rejected.

The encoding used here pads to a fixed width and appends a non-zero terminal digit. Fixed width makes the values sort in the same order as the positions they encode; the terminal digit keeps them legal. Because every value has the same width, appending it cannot change the relative order.

position 1 of 3    →  0.11
position 2 of 3    →  0.21
position 1 of 12   →  0.011
position 10 of 12  →  0.101
Supply sourceOrder in the feed and let the pipeline encode it. Supply orderHint directly only if the source genuinely holds a valid (0,1) string.

If the source has no ordering at all, siblings are ordered by code and a decision is recorded saying so — navigation order is a merchandising decision, not a migration one.

Order hints are scoped to a sibling group, so two categories under different parents legitimately share a value.

Modelling questions worth raising

Is the tree navigation, or facets? A taxonomy used for filtering is usually better as attributes. Flattening a facet-shaped tree into categories produces a navigation nobody wants to browse.
Are there cross-cutting editorial groupings? Seasonal or campaign collections often deserve a second category tree rather than being flattened into a set of codes, which loses their date windows.
Do categories carry attributes? Category-level fields need a Type definition for custom fields. Out of scope for this pipeline today.
Empty leaves. A leaf category with no products and no children renders as an empty page. Reported as a warning, since it is usually a sign the source tree was migrated more completely than the products were.

Other fields

  • externalId preserves the source identifier alongside the key, so downstream systems can still join on it. Populated from the feed's externalId, falling back to the source code.
  • description passes through when supplied.
  • Meta title, description and keywords are not mapped today — they are SEO content, usually owned elsewhere, and inventing them from a name is worse than leaving them empty.
decision-log.md

The engagement decision log

This skill's stated deliverable is a reviewed set of decisions, not a loaded dataset. That needs somewhere durable to live.

Two logs, and why they cannot be one

The pipeline already writes a decision log: out/decisions.json and out/MODEL-REVIEW.md, produced by derive and plan from their own mapping choices. That one is derived — regenerated from the feed on every run, byte-identical for the same input, and gitignored along with the rest of out/. A three-product fixture emits fifteen entries; a real catalog emits thousands.
The log this page describes is accumulated — it records what happened, in order, and must never be regenerated or it loses the only thing it has.
out/decisions.jsonDECISIONS.md
Authorthe pipelineyou, with the user
Naturederived, regenerated each runappend-only, never rewritten
Answerswhat the mapping iswhat was decided, by whom, and why
Scalethousands of entriestens
Lifetimeoverwritten by the next planthe engagement's record

Mixing them fails both ways: a regenerated file cannot hold history, and an accumulated file cannot be trusted to describe the current plan.

Do not copy pipeline decisions into DECISIONS.md. Reference them. One entry recording "reviewed and accepted the 4 irreversible attribute constraints in MODEL-REVIEW.md, plan of 2026-09-22" is worth more than four hundred copied rows that will disagree with the next run.

Where it lives — and the layout it belongs to

"Next to migration.config.json" was ambiguous, because the only config that ships is inside the pipeline, which is the tool — and the tool is now its own repository, cloned rather than authored. Putting an engagement's record there files it inside somebody else's checkout. So be explicit: the engagement is a sibling of the pipeline, never inside it.
<engagement-root>/
  .gitignore                    must ignore the pipeline checkout — see below
  ct-catalog-migration-pipeline/  the tool, cloned and gitignored. Never write
                                  engagement files here, and never edit it — local
                                  changes are lost on the next pull and are
                                  invisible to everyone else.
  source-export/                what the customer handed over
  migration/                    the engagement
    migration.config.json         written from the step-0 interview
    DECISIONS.md                  this log — append-only, committed
    MAPPING-PROPOSAL.md           the step-0b artefact
    adapter/                      the only code written per engagement
      package.json                  {"type":"module"} — see below
    feed/                         generated NDJSON; regenerated, disposable
    out/                          pipeline artefacts; regenerated, gitignored
The adapter needs its own package.json containing at least {"type":"module"}. It sits outside the pipeline checkout, so it does not inherit the pipeline's module type, and an ESM adapter without it fails on the first import with Cannot use import statement outside a module. Thirty seconds to fix and the very first thing an adapter author hits:
{ "type": "module", "private": true }
Commit migration.config.json, DECISIONS.md, MAPPING-PROPOSAL.md and adapter/. Do not commit feed/ or out/ — both are regenerated, and out/ is gitignored by the pipeline already.
The pipeline checkout has to be ignored, and before it is cloned. The engagement is usually a git repository, and a clone inside one becomes an embedded repository that git add -A records as a gitlink — pinned in your index, unexplained by any .gitmodules, and empty for anyone who clones the engagement later. So the engagement's own .gitignore starts as:
/ct-catalog-migration-pipeline/
feed/
out/
The last two are regenerated on every run. The first is somebody else's repository, and its version belongs in DECISIONS.md as a line a human can read rather than as a gitlink nobody can resolve. references/running-the-pipeline.md has the recovery if it was already staged — .gitignore alone does not undo that.
Run the commands from the pipeline checkout and point --config at the engagement:
cd ct-catalog-migration-pipeline
npm run pipeline -- plan --config ../migration/migration.config.json
--out defaults to out relative to the config, the same rule feed.dir follows, so that writes ../migration/out. Do not assume it is relative to the current directory — that reading silently files the artefacts inside the tool's own out/, where the next reader will not look for them.

Format

Numbered, so entries can reference each other and be superseded rather than edited:

# Migration decision log — acme-eu

Append-only. Newest last. Never regenerate; supersede instead.

## D001 — Catalog model: Classic
2026-09-22 · step 0a · decided by Priya (commerce lead)

The largest product has 14 variants, well inside Classic's 100.

Rejected: Modular. It would force standalone pricing and cannot import
`defaultVariant`, and nothing about this catalog needs the higher ceiling.

Status: accepted

## D002 — Colour is an lenum, not free text
2026-09-22 · step 0b · proposed from sample, confirmed by Priya

The sample has 11 distinct colour codes with German and English labels, so
`derive` proposed `lenum`. See `out/probe/MODEL-REVIEW.md`.

**Irreversible.** `changeAttributeConstraint` accepts only `None`, so if the
full export turns out to hold hundreds of colours this becomes an attribute
that has to be recreated and its data rewritten.

Status: accepted, pending confirmation against the full export
Four fields carry the weight: the date, who decided (a decision with no owner is a guess), what was rejected and why (the alternative is what a later reader needs), and status. Mark irreversible and lossy explicitly — they are the entries that will be read under pressure.

What to log, per step

Step 0 — the interview

The config records what was chosen. It cannot record who chose it, what the alternative would have cost, or that a question was asked at all. Log one entry per decision in 0a and 0c, plus the resolved catalog-model/variant-count pair.
If a question could not be answered, log that too, as status: open. An open question with an owner is a plan; an unasked one is a surprise later.

Step 0b — discovery

The mapping proposal is its own artefact. What belongs here is the decisions arising from it: which inferred types were accepted, which were overridden and on what evidence, which source fields are deliberately not migrating, and which findings are hypotheses awaiting the full export.

Steps 1–2 — implementing the adapter

The richest source, and the one most often lost, because these decisions are made while writing code and feel like implementation details. They are not. Log every point where the adapter decides rather than translates:
  • a code derived from a display name, and the derivation rule
  • a hierarchy level collapsed, and what happened to the intermediate
  • an inherited price resolved, and from where
  • a locale tag normalised (en_GB → en-GB)
  • a sentinel or placeholder value skipped, passed through, or interpreted
  • a field dropped, and why
  • every silent cap — anything sampled, truncated or skipped

The adapter reference says a derived code "is a decision, not a detail" and to "report every silent cap". This is where those reports go.

validate and audit

Not every diagnostic is a decision. A finding that changes the adapter is a fix, and fixes do not belong here — the adapter is regenerated and the finding disappears.
What belongs here is a finding that is accepted rather than fixed: 14 category slugs disambiguated and the resulting URLs accepted; a warning about products in no category acknowledged as correct for this catalog; an approaching-variant-limit warning accepted with a plan to revisit. Record the count and the artefact, not the individual findings.

preflight — and this one is an obligation

preflight --apply writes to the customer's project. So does any manual change made to get past it. Those need a durable record, because the next person to look at that project will find settings nobody can account for.
Log, every time: the project key, the timestamp, exactly what changed, and who authorised it. --apply reports project-updated to stdout and nowhere else, so if it is not logged it did not happen as far as anyone later is concerned.
Manual project changes count doubly. Adding a country to make country-scoped prices load is a commerce decision affecting shipping and tax — preflight deliberately refuses to make it, which is exactly why the human who did make it should be named.

load and verify

One entry per --execute: project key, timestamp, what was sent, and the outcome. out/load-result.json has the detail but is regenerated and gitignored.
Log the verify result as the closing entry for a run — including a clean one. "Reconciled, no differences, 2026-09-22" is the sentence someone needs when asking months later whether the load was ever checked.

What not to log

  • Pipeline mapping decisions, copied. Reference the artefact.
  • Fixes. A defect found and corrected leaves no decision behind.
  • Anything the config already states. Log the why, not the value.
  • Conversation. An entry is a decision with an owner and a rationale, not a transcript.

The discipline that makes it worth keeping

Write the entry when the decision is made, not at the end. A log reconstructed after the fact records what someone remembers deciding, which is reliably the subset that turned out well — and the entries that matter are the irreversible ones taken under uncertainty, which are the first to be forgotten.

pricing-and-money.md

Money and prices

The two things most likely to be wrong in a way that looks right.

Minor units

commercetools money is expressed in minor units (centAmount), and the digit count is not 2 everywhere:
Fraction digitsCurrencies
0JPY, KRW, and others
2most
3BHD, KWD, OMR, TND
Defaulting to 2 does not error. It multiplies a 0-digit price by 100 and loads it. market.currencyFractionDigits therefore requires an explicit entry per currency, and a missing one is a hard error rather than a default.

Convert on the digit string, never through a float

A binary float cannot represent most decimal amounts exactly. 0.07 * 100 is 7.000000000000001. The error surfaces as an off-by-one cent on a fraction of records — the worst kind of migration defect, because everything looks fine.

Amounts travel through the feed as decimal strings and are converted by string manipulation:

"19.99"  GBP (2)  →  1999
"4500"   JPY (0)  →  4500        not 450000
"12.345" BHD (3)  →  12345
"0.07"   EUR (2)  →  7
Excess precision is refused, not rounded. Rounding would silently change a price, and nobody reviews a price they were not told changed:
"29.999"  GBP (2)  →  error: 3 decimal places but GBP allows 2
"4500.5"  JPY (0)  →  error
Trailing zeros beyond the currency's precision are not a loss. "29.9900" is exactly 29.99 and is accepted.
Verify after mapping: plan --payloads writes every price decoded back from minor units to a decimal, so the numbers can be checked against the source.

Price scope

A price's scope is currency, country, customer group, channel, and validity window. Only one price may exist per scope.
A store is not part of that scope. There is no store field on a price, and there never was — a store scopes prices indirectly, by listing the distribution channels it trades through. In a store's context the candidate prices are the ones whose channel it lists, plus every price with no channel. The channel is the pricing unit; the store is only the binding. Two distribution channels on one store, each with a price for the variant, match at the same fallback step and the storefront has to pick with setLineItemDistributionChannel. See catalog-feed-contract.md for the store record.
Two of those are references to resources the Import API cannot create: it has no channel or customer-group type. So a price scoped to either must declare it in the feed (channel / customerGroup records), and load creates the missing ones through the platform API before any import runs. Without that, the price becomes an operation that sits unresolved for 48 hours and then expires — a green load and a missing price.
validate refuses an undeclared reference. preflight reports one the project is missing as a warning, since load will create it, and one that exists without the roles the plan needs as an error, since that is the case load refuses to fix. See catalog-feed-contract.md.

Adding an embedded price is rejected when:

  • another embedded price has the same scope and the same validFrom / validUntil, or
  • two prices have overlapping validity periods within the same scope.
One nuance that matters, and that a naive check gets wrong: a price with no validity period does not conflict with a price defined for a time period. A base price plus a dated promotion in the same currency and country is the normal shape, and flagging it would be a false positive that gets the gate switched off.

Validity intervals are half-open, so one window ending exactly where the next begins does not overlap.

The audit gate checks all of this offline. Resolving inherited prices produces duplicate-scope shapes very easily, which is why it is worth catching before the API names one offending record and stops.

Field name

The feed calls it validTo. The API field is validUntil. The mapper translates; hand-written payloads often do not.

Embedded versus standalone

EmbeddedStandalone
Storedinside the Product Variantas independent resources
Limit100 per variant (soft)50,000 per variant (soft)
Catalog modelClassic onlyboth — and the only option under Modular
Scope neededmanage_productsmanage_standalone_prices
Scoped price search (filter/facet/sort)supportedinconsistent — only embedded prices are considered
Both modes work on Classic. Modular allows only standalone — it has no embedded prices at all, a VariantImport has no price field, and the config loader refuses the other pair. priceMode: standalone emits StandalonePriceImport resources keyed by SKU as the final load stage, after the products (and, under Modular, after the variants those prices reference by SKU).
Both limits are soft — increasable per Project after a performance review, like the category and ProductType counts. See limit increase guidance. Treat them as a modelling signal rather than a wall: hitting 100 embedded prices on a variant usually means the price scopes want rethinking, not that the limit wants raising.

Mixing both types on one product is possible but degrades performance, so the pipeline takes a single choice.

Three things about standalone pricing that only bite at load time:

  • It needs manage_standalone_prices. manage_products does not grant it, unlike every other import request this pipeline sends. A client set up for an embedded load fails on that one stage with a 403 and nothing else.
  • sku is a plain string, not a KeyReference, and the API does not validate it. A price for a SKU no variant has is created successfully and prices nothing. The audit gate checks every SKU against the plan's own variants, because that is the only place it can be caught.
  • priceMode on the product decides which prices are read. A product declaring Embedded whose prices are standalone resources imports with every operation reporting imported, and then shows no price anywhere. The gate treats any disagreement between the config, the products and where the prices actually are as an error.

Uniqueness differs between the two, and so does what a duplicate means:

EmbeddedStandalone
Unique pervariant + scopeSKU + scope (project-wide)
Overlapping validity in one scoperejectedaccepted

Because standalone overlaps are accepted, the gate reports them as a warning rather than an error: the API will take them, but which price wins stops being something the migration decides. An open-ended base price alongside a dated promotion is not an overlap in either mode — that pair is the normal shape of a priced catalog.

Keep the embedded price count low by leaning on price-selection fallback: one price for a currency with no country beats one price per country, and a channel-less base price beats a price per channel.

Things with no embedded-price equivalent

Report these as information loss rather than approximating them:

  • Quantity breaks / tiered pricing. Native tiers exist on Standalone Prices; an embedded price has no tiered equivalent. A quantity break usually becomes a Cart Discount, which is a modelling decision, not a translation.
  • An unmappable customer-group price. Loading a trade or staff price unscoped shows it to every shopper and collides with the base price as a duplicate scope. Skip and report. Skipping is recoverable; wrong data in production is not.
  • Net versus gross. Whether a price includes tax is usually a setting elsewhere in the source, not a property of the price row, so a price read on its own does not tell you what its number means. Getting it backwards changes every price on the site by the tax rate and nothing in the data contradicts you. It maps to a tax mode, which is a modelling decision.
  • A tax code is not a tax rate. A product tax code is a reference into an external tax engine. Carry it across verbatim; never derive a rate from it.

Currency and locale acceptance

A project rejects money in a currency it does not accept, and localized strings in locales it does not accept. The audit gate checks every price currency against market.requiredCurrencies, and making the project accept them is a prerequisite of the load — an idempotent, additive step, not a manual click.
product-model.md

The product model

How ProductTypes and their attributes are derived, and which of those choices cannot be undone.

Reference: Product Types, attribute types. Fetch the api-ProductType schema from the commercetools Knowledge MCP when you need exact field shapes.

Declared or inferred

Declared — the feed carries attributeDefinition records because the source has a type system. Derivation is a translation, and every choice has a defensible origin.
Inferred — the feed carries none, so definitions are reconstructed from observed values. Every guess is recorded with review: true and has to be signed off. Inference is a fallback; emitting declarations from the adapter is the reliable fix.
Inference reads values, never field names, and refuses to guess rather than guessing wrongly: an attribute holding both strings and numbers is an error, not a coerced text.

What inference can and cannot recover, on the same catalog:

DeclaredInferred
Types, levels, constraintsfrom the declarationreproduced from values
A localized enumlenum with labelsdegrades to enum, code as label
Labelsas supplied if the source has any — many type systems have only a description field, and then declared attributes get humanized names toohumanized attribute name, default locale only
Unitscarried into the labellost — nothing to carry
isRequiredhonoured when fully populatednever inferred
Set productTypes.onMissingDefinitions to require to refuse to guess at all.
Declaring types does not guarantee labels. The table above reads as though declaration buys readable labels; it buys them only where the source carries them. A hybris items.xml has <description> — developer documentation, not a merchandiser-facing label — and no label field at all, so a fully declared type system can still see the large majority of its attributes fall back to humanized names. Check whether the source has a real label field during discovery, and if it does not, say so in the mapping proposal: the Merchant Center will show Fabric composition derived from fabricComposition, which is usually acceptable and occasionally not.

Attribute constraints are irreversible

changeAttributeConstraint accepts only None — AttributeConstraintEnumDraft has exactly that one value. A constraint can therefore be relaxed later, but never tightened or switched.
Getting it wrong means recreating the attribute and rewriting its data. This is the single most consequential decision in the whole migration, which is why derive records every non-None constraint as irreversible and MODEL-REVIEW.md leads with them.
Source situationConstraintWhy
Part of variant identityCombinationUniquethe platform then enforces one variant per combination
Invariant across a product's variantsSameForAllvariants cannot disagree
Varies freely per variantNoneno constraint to enforce
Must differ on every variantUniquerarely what a migration wants
SameForAll enforces that variants agree; it does not distribute a value. A product-level value therefore has to be written onto every variant, and an attribute present on some variants but not others is refused.

Product level versus variant level

AttributeLevelEnum is Product or Variant, and the trade-off is real:
productTypes.productLevelStrategy defaults to sameForAll for that reason. Choose native only when the storefront is known to use Product Search exclusively.
Do not repeat keys.prefix in defaultKey. Every resource key is prefixed, ProductTypes included, so keys.prefix: "acme" with defaultKey: "acme-apparel" produces acme-acme-apparel. validate warns (product-type-key-doubles-prefix); the key cannot be changed once products reference the ProductType.
The names read backwards, so read them twice. sameForAll is the safe default and native is the one that makes attributes invisible to Product Projection Search — so the cautious-sounding option is the risky one, and the option that sounds like a workaround is the one to pick. The values name what the pipeline does (write a variant attribute constrained SameForAll, or use the platform's native Product level), not how safe they are.

An attribute name is a Project-wide declaration

Two invariants hang on this, and neither is scoped to a ProductType.

The type must be identical. An attribute name may hold exactly one type across the whole Project. Two ProductTypes that share a name must agree, and the API enforces it with AttributeDefinitionTypeConflict — after which every product referencing that attribute fails AttributeNameDoesNotExist. So a plan that is internally consistent can still be refused by a ProductType the migration never touches:
A load declared material as text. An unrelated existing ProductType had it as ltext. Both operations were rejected.
derive catches the clash within a plan; preflight is the only stage that can catch it against the project, because it is the only one that reads the project. Enum values may differ freely — the documentation's own example gives Color a different value set on Jeans than on T-Shirt. Only the type is constrained, along with a set's element type and a reference's target.

There is no update action that changes an attribute's type. On a ProductType that already exists, the fix is to remove and re-add the attribute, which deletes its values on every product using it. Treat it as a migration decision.

isSearchable must agree too. To use one attribute name across several ProductTypes for search, filters or facets, isSearchable must be true on all of them. If the values differ, the attribute becomes unavailable for search, filters and facets everywhere.
That one fails silently: the import succeeds and the facet is simply missing. derive treats a mismatch within the plan as an error, and preflight warns when the plan disagrees with the project — which can break a facet that works today.
Axes are always searchable. Everything else follows productTypes.searchableByDefault.

Grouping products into ProductTypes

Products group by their declared productType, falling back to the configured default. A ProductType's attribute set is the union across its members, since any member may set any of them.

One ProductType per product is a modelling mistake — it defeats the point of a shared blueprint and burns through the project limit of 1000 ProductTypes — so the grouping never invents keys.

Three clashes are fatal, because a ProductType cannot express them:

  • An attribute that is an axis for some products in the group and a plain attribute for others. Constraints belong to the ProductType, not the product. Split the ProductType, or make the axis consistent.
  • An attribute at product level on some records and variant level on others. One attribute cannot be both SameForAll and free-varying.
  • An attribute holding more than one kind of value. Declare it explicitly or fix the adapter.

Choosing types

Source shapeTypeNote
Plain stringtext
Per-locale textltext
Fixed value set with stable codesenum / lenumkeys must be codes, not display text
Numbernumberno unit concept — see below
Booleanboolean
Calendar date, timestamp, timedate, datetime, timeISO form; inferred only on full agreement
Multiple valuesset of the element type
A unit is appended to the label, whatever the type. commercetools attributes have no unit concept anywhere, so unit on an attributeDefinition is carried into the label and is no longer machine-readable — a consumer cannot convert or compare across units. Recorded as information loss.
This is not limited to number. A measurement the source expresses as a range ("18–24 °C") has no single numeric type to go in, and there are two honest mappings rather than one right answer:
KeepsLoses
One text / ltext attributethe source's exact wording, including any "up to" or "approx."comparability — and facetability: with searchableByDefault: true a value like "20 - 30" becomes a filter nobody can use
Two number attributes (…Min, …Max)filtering, sorting and range queriesthe original phrasing, and any range that is not two plain numbers
Pick on whether the attribute has to be searchable. That is the deciding factor and it is easy to miss: a faithful text range is the safer-looking choice and quietly produces a useless facet. Splitting is a decision worth logging either way — the source said one thing and the project now holds two.
unit behaves the same in both: it reaches the label and no further.
Date, time and datetime have exact accepted forms, checked by the audit gate and used by inference:
TypeFormExample
dateYYYY-MM-DD2026-09-30
timeHH:MM, optionally :SS or :SS.sss14:30, 14:30:00.000
datetimeYYYY-MM-DDTHH:MM plus seconds, and an offset or Z2026-09-30T00:00:00Z
A datetime with no zone is rejected. That matters when the source has bare dates: 30/09/2026 becomes 2026-09-30T00:00:00Z, which is midnight at the start of the day, so a window ending on it excludes that day. Whether the source meant inclusive is a question for the customer, not a coercion rule — decide it, log it, do not silently add a day.
A value's own label is separate from the attribute's. The fallback rules above are about the attribute label. Each enum / lenum value carries its own, and an omitted one falls back to the key in the default locale:
Declared asValue label suppliedResult
lenum{"en-GB": "Warm"}used as given
lenum"Warm" (plain string){"<defaultLocale>": "Warm"}
lenumomitted{"<defaultLocale>": "<key>"}, recorded as lossy
enum"Warm"used as given
enum{"en-GB": "Warm", "de-DE": "Warm"}the default locale's value; others dropped, recorded
enumomittedthe key, recorded as lossy
Two consequences worth planning around. A plain enum holds one label per value, so a source with per-locale display text needs lenum or the other locales are lost. And a derived label means merchandisers see the raw code in the Merchant Center — small, but it is loss, so it appears in MODEL-REVIEW.md rather than happening quietly.
Low-cardinality strings are not automatically enums. Enum keys must be stable codes, and source strings are usually display text. Inference maps them as text and raises the enum question rather than minting unstable keys.
Nested and set-of-nested attributes are not searchable and cannot target discount predicates. Do not choose them for anything that must facet or drive a promotion. A set chain terminating in nested is limited to 5 steps.

Review the source model; do not just translate it

The most valuable output of a migration is often a defect found in the source. Worth raising rather than silently carrying across:

  • A variant axis with no stable code, only a localized name.
  • Categories that are really facets.
  • Attributes that encode business rules.
  • Attributes declared but never populated — usually a field the adapter dropped, and reported as such.
  • A 300-attribute ProductType serving eight facets.

Output

derive writes:
  • out/product-types.json — the drafts, for the load stage and for diffing.
  • out/decisions.json — the machine-readable decision log.
  • out/MODEL-REVIEW.md — the sign-off artefact, grouped into irreversible choices, information loss, and guesses. Nobody reviews a JSON file.

Every decision carries a rationale. The rationale is not a comment; it is the deliverable. A mapping added without one is an incomplete change, because the next reader cannot tell whether the choice was considered or accidental.

running-the-pipeline.md

Running the pipeline

Seven stages, each gating the next. The first four need no credentials at all: a defect is reported at the earliest stage that can see it, so a broken feed or an unloadable plan says so on a laptop with no .env.

Getting the pipeline

The pipeline is a separate repository, cloned beside the engagement rather than into it. It is a tool, not engagement content: nothing in it is edited per engagement, and a teardown scoped to keys.prefix has nothing to do with it.
Ignore it before you clone it. The engagement is usually itself a git repository, and a clone inside one is an embedded repository: git add -A stages it as a gitlink — mode 160000, the pipeline's commit pinned in your index, no .gitmodules to explain it. That looks tracked and is not. Anyone who clones the engagement afterwards gets an empty directory and no way to obtain the contents:
warning: adding embedded git repository: ct-catalog-migration-pipeline
hint: Clones of the outer repository will not contain the contents of
hint: the embedded repository and will not know how to obtain it.

So the order matters:

# from the engagement root, BEFORE cloning
echo '/ct-catalog-migration-pipeline/' >> .gitignore

git clone https://github.com/commercetools/ct-catalog-migration-pipeline
cd ct-catalog-migration-pipeline
npm install            # first time; Node 20+
npm test               # no network, no credentials — a few seconds
If it is already staged, .gitignore will not undo it — ignore rules do not apply to anything already in the index, and git rm --cached refuses without -f because the gitlink differs from both the worktree and HEAD:
git rm -r --cached -f ct-catalog-migration-pipeline
A submodule would be the other way to do this, and is the wrong one here: it pins the version in the outer repository's history, which is what the DECISIONS.md line below already does in a form a human can read — without making every clone of the engagement fetch a second repository.

Run the tests even when cloning a known-good commit. "The four stages ran clean on my feed" is the tempting substitute and a much weaker signal: a pipeline that works on one feed says nothing about whether the checkout is intact. A few seconds against a stage misbehaving later for a reason you cannot localise.

Record which commit you used in DECISIONS.md. This matters more than it looks: the pages here name exact diagnostic codes, field names and behaviours, and the skill and the pipeline are versioned separately. A pairing that has drifted produces documentation that describes a tool you are not running — which is the single most common defect class this project has had.
These pages were checked against main on 2026-09-24. If a stage, flag, diagnostic code or fixture named here is absent from your checkout, trust the checkout and treat this page as the stale half of the pair.

One line is enough:

Pipeline: ct-catalog-migration-pipeline @ <short sha> (git rev-parse --short HEAD)

The stages

npm run pipeline -- validate --config ../migration/migration.config.json
npm run pipeline -- derive   --config ../migration/migration.config.json
npm run pipeline -- plan     --config ../migration/migration.config.json --payloads
npm run pipeline -- audit    --config ../migration/migration.config.json
Common options: --config <path>, --out <dir>, --json for machine-readable diagnostics, --quiet to suppress warnings, and --payloads on plan.
--out defaults to out resolved relative to the config file, the same rule feed.dir follows — so running from the pipeline checkout with --config ../migration/migration.config.json writes to ../migration/out, not into the tool's own directory. An absolute --out is used as given. The engagement layout this assumes is in decision-log.md.

Every stage exits non-zero when it finds an error, so the sequence composes in a shell or CI.

What is built

StageNetworkStatus
validate — feed against the contract, and against the confignobuilt
derive — ProductTypes, declared or inferrednobuilt
plan — map to drafts, write the decision lognobuilt
audit — every write-time invariantnobuilt
preflight — project catalog model, locales, currenciesyesbuilt
load — Import APIyesbuilt
verify — read back and reconcile (GET only)yesbuilt
Every stage is built and tested. load is a dry run unless --execute is passed, so nothing writes catalog data by accident, and verify cannot write at all.
Treat a green audit, preflight and load as "the plan is internally consistent, the project accepts it, and the operations were submitted" — not as "the catalog is right". Those are different claims, and load cannot make the second one: the Import API validates asynchronously, so every request can come back accepted while individual operations are rejected afterwards. verify is what closes that gap. Run it.
Still absent by design: a cutover sequence and a delta strategy, which are engagement decisions rather than pipeline features. verify also does not yet scan for orphans across the key prefix — it checks what the plan names, not what the project holds that the plan has forgotten. Say that out loud rather than letting green stages imply readiness.

validate

Schema conformance, then catalog-wide integrity — and integrity only runs once every record parses, because a rejected record drops a product from the index and makes the integrity pass invent orphans that are really just consequences of the earlier failure.

Catches: malformed JSON, schema violations, unknown record types, duplicate codes and SKUs, missing category parents, category cycles, orphan variants, products without variants, missing category references, axis coverage and uniqueness, localized axes, axes declared at product level, attributes written but not declared, level mismatches, multiple master-variant claims.

It also checks the feed against the config, which is not a contract concern but is something only this stage can see early:
  • A price scoped to an undeclared channel or customer group. An error: load creates the ones the feed declares, so an undeclared reference is one nothing will create, and it expires unresolved after 48 hours taking the price with it. A price channel also needs the ProductDistribution role — an error under standalone pricing, which the API rejects outright, a warning under embedded.
  • A dangling or self-defeating assortment. A product assigned to an undeclared selection (the assignment is silently dropped — nothing fails, the assortment is just wrong); a store listing an undeclared channel, or one whose roles do not match the list it is in; an empty Individual selection, which offers nothing; and a store whose selections are all inactive with at least one Individual, which offers no products at all where an empty list would have offered everything.
  • Relative image or asset-source URLs with no media.baseUrl. An error, not a warning: the host is not in the export, a wrong one 404s every image undetectably, and the contract previously forced authors to invent one by rejecting relative paths outright.
  • A price in a currency market.requiredCurrencies does not list, reported once per currency with the first price that uses it. Left to plan it would surface once per affected variant, two stages later.
  • Which catalog model the catalog requires. Classic caps at 100 variants per product, Modular at 10,000, and that is purely a function of variant counts — so the feed can settle it with no credentials and no plan. The report states the largest product and the model it needs even when Classic fits, because "Classic is enough" is the answer someone is looking for.
A catalog needing Modular while the config says Classic is a hard error, because the import shape is fixed at map time: the remedy is to set target.catalogModel and re-run plan, not to edit the config and load anyway. The stage footer drops its usual "fix it in the adapter" advice when the catalog model is the only error — the feed is correct, and splitting products to fit Classic would be reshaping a catalog to suit the tool.

Diagnostics name a file and a line. Fix the adapter, not the feed.

derive

Builds ProductTypes from declarations, or infers them from observed values. See product-model.md.
Reading out/MODEL-REVIEW.md is the point of this stage. It groups what needs a human into irreversible choices, information loss needing sign-off, and guesses.
Fatal here: an attribute that is an axis for some products in a ProductType and a plain attribute for others; one at product level on some records and variant level on others; one holding mixed value kinds; and isSearchable disagreeing across ProductTypes for a shared name.

plan

Constructs the drafts. It does not verify — that is audit's job, and keeping them apart is what makes the audit trustworthy.
Four things are resolved here, each covered in its own reference: minor-unit money conversion, slug allocation and disambiguation, order-hint encoding, and deterministic master-variant selection.

Two structural points worth knowing when reading a plan:

  • Product-level attributes appear on every variant, because SameForAll enforces that variants agree rather than distributing a value.
  • Axis values appear as ordinary attribute values holding the enum key. axisLabels never reach a variant.

Output:

FilePurpose
out/plan.jsonwhat the load stage consumes
out/key-map.jsonsource identifier → key, for delta runs and rollback
out/decisions.jsonthe full decision log
out/sample-payloads.mdwith --payloads: one product per variant shape, prices decoded back to decimal
Products are always planned with publish: false. Import staged, review in the Merchant Center, publish deliberately.
On a first load that is what you want. On a re-load into a project whose products are already published, it takes the catalog offline. publish is not "leave publication alone" — it is an instruction, and false is an instruction to unpublish. Per the Import API best practices:
publishImport carries changesResult on an existing Product
falsenothe Product is unpublished
falseyeschanges go to staged, the Product is unpublished, hasStagedChanges becomes true
trueyesapplied to current and staged. ProductDraftImport publishes an unpublished Product; ProductImport does not
truenostaged changes are applied and the Product is published
Nothing errors. Every operation reports imported, and the storefront — which reads current — goes blank for every product the run touched. So before any --execute against a project that is already live, establish whether its products are published, and treat re-loading a live catalog as its own decision with its own entry in DECISIONS.md. A first load into an empty project is the case this default was designed for; it is not the only case a pipeline gets pointed at.

Plans are deterministic: the same feed produces a byte-identical plan, which is what makes a diff between runs meaningful.

audit

The gate. It reads plan.json back from disk and re-derives every invariant from the drafts as data — it never touches the feed or the mapper.

That independence is the point. A verifier sharing logic with the writer confirms the writer's bugs rather than finding them, and reading the file means the gate checks what will actually be sent. It also keeps working on a hand-edited or replayed plan.

Reading from disk has one cost, and it is checked

The plan and the feed can drift apart, and the shape is this: plan fails, writes nothing, and leaves the previous run's plan.json in place. audit then reports zero errors — against a plan that no longer describes the feed, which by now contains a duplicate SKU. That is the exact invariant the gate exists to catch, and the gate cannot see it: the stale plan is internally consistent, just describing data nobody is going to load.
So plan stamps plan.json with provenance.feedDigest, a SHA-256 over the feed's *.ndjson files, and both audit and load refuse a mismatch:
StateCodeSeverity
Digest matches the feed—nothing reported
Digest differsplan-staleerror — nothing audited, nothing loaded
No digest at allplan-provenance-unknownwarning — re-run plan to stamp it
An absent digest is deliberately not an error: a hand-written plan has none, and the gate is meant to work on those. But it cannot be called fresh either, so it reports as unknown rather than passing quietly.
The digest covers file names as well as contents — a rename changes what the plan claims provenance from — but not the directory path, so moving an engagement or reading it through a different relative path keeps the plan valid.
Thirty-four codes block the load. The count was stale here for a while, and a list that has stopped matching the code is worse than no list — this one is generated from severity: 'error' in audit/gate.ts:
AreaCodes
Identity and content (10)duplicate-resource-key, duplicate-sku (project-wide), duplicate-slug (per locale), duplicate-price-key, variant-without-sku, invalid-key, invalid-slug, invalid-order-hint (outside (0,1) or ending in 0), missing-name, dangling-product-type
References (2)dangling-category-parent, dangling-category-reference
Attributes (5)attribute-not-declared, attribute-type-mismatch (including set element types), enum-value-not-declared, duplicate-attribute, required-attribute-missing
Constraints (2)same-for-all-violation (values disagreeing, or present on only some variants), combination-unique-violation
Prices (4)duplicate-price-scope, overlapping-price-validity, currency-not-configured, fraction-digits-mismatch
Price mode (5)product-price-mode-mismatch, embedded-prices-in-standalone-mode, standalone-prices-in-embedded-mode, duplicate-standalone-price-scope, standalone-price-orphan
Assets (4)asset-without-source, asset-without-name, duplicate-asset-key (per variant or category, not per project), duplicate-asset-source-key (within one asset)
Limits (2)variant-limit-exceeded (100 per Product under Classic, 10,000 under Modular), price-limit-exceeded (100 embedded per Variant, embedded mode only)
Eight more are warnings — they describe a catalog that will load and may still be wrong: approaching-variant-limit, attribute-never-populated, category-without-products, no-prices-planned, order-hint-absent, overlapping-standalone-price-validity, product-without-category, variant-without-price.
Six more are advisory: attribute-never-populated, variant-without-price, product-without-category, category-without-products, order-hint-absent, and approaching-variant-limit (past 80 variants, before the cap bites).
Findings are grouped by check and capped per code, because one broken assumption in an adapter produces hundreds of instances of a single code and a flat list buries every other finding. Use --json for the complete set.

Cascades are suppressed: a product whose ProductType does not exist is not also reported for every missing required attribute.

preflight

A handful of GETs that answer one question: will this project accept this plan? Four always — project settings, and counts of product types, categories and products — plus one each for channels and customer groups when the feed declares any, and a paginated read of every ProductType with its attributes. Every failure it catches would otherwise surface as a wall of rejected Import Operations, hours into a load, one error per record.

cp .env.example .env
npm run pipeline -- preflight --config ../migration/migration.config.json
npm run pipeline -- preflight --config ../migration/migration.config.json --apply

It prints the project it is about to write to first, with existing resource counts, because loading into the wrong project is a real hazard.

Blocking: the project is unreachable or the credentials fail; productCatalogModel disagrees with target.catalogModel; the project does not accept a locale or currency the plan needs; the default locale is not accepted; a price country the project does not list, which the API rejects outright; a declared channel that exists without a role the plan needs, since that is the one thing load will not fix for itself; an attribute name the plan defines with a different type from the one the project already has for that name. Advisory: a non-empty project, counts that could not be read (a client with project-settings scope but no product scope still gets a useful preflight), a channel or customer group that is simply missing — that one is a warning because load creates it — an attribute name whose type matches but whose isSearchable does not, a project that still holds ProductTypes under the plan's key without keys.prefix, a product selection or store that is simply missing (both get created), and — as errors — a product selection whose mode already differs, or a store whose wiring already differs, or a store declaring a language or country the project does not accept.
A selection's mode is the sharpest of those: there is no changeMode action, so importing over an existing selection cannot apply the plan, and the assortment ends up meaning the opposite of what was intended.
It checks what the plan needs, not only what the config declares — a plan can carry a locale nobody remembered to list.
--apply adds missing locales and currencies, and nothing else. It does not create channels or customer groups: --apply changes project settings, and creating resources is load's job, where it is reported alongside everything else that run wrote. The update sends the union of existing and needed values, because changeLanguages and changeCurrencies replace the whole array: sending only the missing entries would delete every locale and currency already configured. A 409 refetches and retries with the fresh version.
It will not change productCatalogModel — that is a decision about the whole implementation — and it will not change countries, which also drives shipping and tax.

A project loaded before ProductType keys were prefixed

ProductType keys used to be emitted verbatim from the feed's productType code (or productTypes.defaultKey), so a mig engagement created mig- everything except its ProductTypes. They are prefixed now, which means a project loaded before the fix holds apparel-basic where the plan says mig-apparel-basic.
preflight reports that as product-type-keys-unprefixed. It is a warning, because on an empty project or one loaded since the change this is simply the normal case — but the consequence when it does apply is sharp: the load creates the prefixed ProductType and then cannot move the existing products onto it, because a Product's ProductType cannot be changed after creation and there is no update action for it.

On a throwaway project, delete the products and re-load. On one holding data worth keeping, this is a migration decision: log it and decide deliberately.

The ProductType's name is unaffected — it still comes from the unprefixed code, so it reads "Apparel (basic)" rather than "Mig Apparel Basic".

The one project-wide constraint only preflight can see

An attribute name may hold exactly one type across the whole Project. Two ProductTypes with nothing else to do with each other must agree, so a plan can be internally consistent and still be refused by what the project already holds.
This is the third project-wide constraint here, after SKU uniqueness and slug uniqueness, and it is the only one that needs the project to be checked at all: audit is offline, so it cannot see it, and the feed cannot either.
The cascade is worth recognising. A project whose apparel-basic holds material as ltext, against a plan declaring it text, returns AttributeDefinitionTypeConflict on the ProductType — and then AttributeNameDoesNotExist on every product that references it. Two rejected operations for one avoidable mismatch, and the second error names a symptom rather than the cause.
  • A differing type is an error. There is no update action that changes an attribute's type: on an existing ProductType it would have to be removed and re-added, deleting the values on every product using it. Preflight says so explicitly when the clash is on a ProductType the plan itself loads, because the remedy is different — a migration decision, not an edit.
  • Enum values may differ freely. The documentation's own example gives Color a different value set on Jeans than on T-Shirt. Only the type is constrained, so preflight compares type names — and, for a set, its element type; for a reference, its target.
  • A matching type with a differing isSearchable is a warning. The load succeeds, and the attribute then becomes unavailable for search, filters and facets across every ProductType — including ones working today. Nothing errors; the facet simply stops existing.
derive enforces the same rule within a plan (type-mismatch-across-product-types), and as an error, not a warning: the API refuses the clash outright. It is not the milder "legal, but a storefront has to handle both shapes" that the two-shapes reading suggests.

load

Dry run by default. --execute is the only thing that writes catalog data.
npm run pipeline -- load --config ../migration/migration.config.json              # dry run
npm run pipeline -- load --config ../migration/migration.config.json --execute
npm run pipeline -- load --config ../migration/migration.config.json --execute --wait
A dry run writes out/load-requests.json with the exact request bodies, and the prerequisites it would create. Read that before executing — it is what will be sent, not a summary.
The audit gate runs again inside load rather than trusting it was run. It is free, and the alternative is finding a rejected invariant one request at a time after part of the catalog has landed.

Two platform stages, at opposite ends of the run

Three stages do not go through the Import API at all, and they do not all run at the same time:

StageWhenWhy
channelfirstprices reference them, and an unresolved price expires after 48 hours
customer-groupfirstsame
storelastit references product selections, which the Import API creates asynchronously
platformStages(order, phase) takes the phase explicitly rather than defaulting, because a caller that forgot it would create stores before the selections they point at exist.
load refuses without --wait when a store references product selections (stores-require-wait). Without waiting there is no moment at which the store stage is safe to run, and stopping before the import is cheaper than a catalog whose stores were never wired.
An existing store is never modified. setDistributionChannels and setProductSelections replace the whole array, so applying the plan over a store the project already configured would discard wiring this migration knows nothing about. Both preflight and load report the difference (store-wiring-differs) and stop.

Prerequisites run first, and not through the Import API

The Import API has no channel or customer-group resource, so load creates those two through the platform API, before the first import request. A price whose channel does not exist yet becomes an operation that expires unresolved after 48 hours, so the ordering is load-bearing rather than tidy.

The rule is read first, then create only what is missing:

  • an absent channel or group is created;
  • an existing one is left untouched — not patched to match the plan;
  • an existing channel lacking a role the plan needs is an error: roles govern stores and inventory as well as prices, so widening them is a project decision, not something a catalog load does in passing;
  • a failed read creates nothing, because creating without knowing what exists is how duplicates happen.

They are reported separately from the containers, because they have no container, no operation states and no 48-hour window — the call either succeeded or it did not.

An unmet prerequisite stops the import before it starts. If one is unusable, could not be created, or could not even be read, load reports prerequisites-unmet and sends no import request and creates no container. Proceeding would import prices scoped to something absent, and those operations expire unresolved — a run reporting every request accepted, over a partly priced catalog, with no signal left by the time anyone looks. Stopping costs nothing a re-run does not recover, because every key is deterministic and a second load updates rather than duplicates.
A dry run performs the reads too, which costs only GETs and is what lets it name the exact keys it would create instead of guessing. If that read fails the dry run still completes, with a warning and those keys reported as existence unknown — an unreachable project should not cost you the request bodies.
Two asymmetries worth knowing: these are the only resources keyed verbatim rather than <prefix>-<key>, and therefore the only ones a prefix-scoped teardown leaves behind. If load created a channel, removing it is manual.

How large a catalog this handles

Not a memory question, and more RAM does not move it. Every artefact is written with JSON.stringify, which builds the whole document as one string, and V8 caps strings at about 537 MB whatever the heap size. Measured on this pipeline:
Catalogplan.json per variantCeiling
2 product-level attributes0.9 KB~575,000 variants
25 product-level attributes4.2 KB~125,000 variants
The spread is productLevelStrategy: sameForAll, the default: a product-level attribute is stated once in the feed and written onto every variant in the plan. Measured amplification, feed to plan: 15× (2.8 MB → 42.6 MB for 10,000 variants). So the more attributes a catalog carries, the sooner it arrives — which is the opposite of the intuition that variant count is what matters.

The feed held in memory is the lesser cost: ~14 KB of RSS per variant, so a 125,000-variant run peaks near 1.8 GB against a 4.3 GB default heap.

Past the ceiling the pipeline says so — naming the artefact, the resource counts, and that RAM is not the problem — rather than emitting V8's bare Invalid string length. The same limit applies on the way in: audit, load and verify each read plan.json as one string, so a plan can be written and then be too large to read back.
The way through today is to split the engagement: several configs with different keys.prefix values, each covering part of the catalog. Each plan is then its own document and the loads stay additive. Writing the artefacts incrementally — NDJSON per collection — is what would remove the ceiling, and it changes an artefact the freshness digest and the dry-run review both depend on, so it is worth doing when an engagement needs it rather than in advance.

Batching

  • At most 20 resources per Import Request, clamped in code.
  • Containers named by resource type (mig-category, mig-product-draft) and reused across runs; a run-scoped name would exhaust the 1000-container limit and hide the previous run's work. Each is restricted to its resourceType, so a mis-routed batch is rejected.
  • A stage beyond the per-container operation limit splits into numbered containers, and no request straddles two.
Reuse has a clock on it. A container created without a retentionPolicy is deleted 72 hours after creation — not 72 hours after last use, so the window does not extend as a migration continues. A TimeToLiveRetentionPolicy set at create time takes a custom duration instead. This bites on any engagement that spans more than three days: the deterministic name resolves to a container that no longer exists, and with it goes the operation history the name was reused to preserve. Keys are deterministic and a re-created container takes the same name, so nothing is lost that a re-run cannot redo — but plan the retention rather than discovering the gap.
Note the two clocks are different and neither resets: operations are deleted at 48 hours, containers at 72. So a container can outlive the operations it holds, and a summary read late in that window reports fewer operations than the run submitted.
What --wait covers. It waits for operations to leave processing — i.e. for the API to finish validating what was sent. It does not wait for unresolved operations to find their KeyReference targets, because that can legitimately take up to 48 hours and depends on data this run may not be sending. Expect unresolved counts after a --wait load of a deep category tree, and expect them to fall on their own.

Do not serialise on polling

Stages are pushed in dependency order but not awaited. An unresolved operation completes on its own when its KeyReference target arrives, any time inside 48 hours, so waiting per stage serialises a pipelined design — and frequent summary polling actively slows the import.
--wait polls with doubling backoff until nothing is processing, and says plainly when it times out.

Which states are failures

StateTreated asWhy
unresolvedwarningcompletes on its own inside 48h
waitForMasterVariantwarningresolves when the master variant lands
processingwarningasynchronous; re-read the summary
rejectederrorthe only resubmittable state; others retry internally up to 5 times
validationFailederrorthe gate should have caught it
partiallyImportederrorthe resource is half-written
A failed request records the keys it carried instead of aborting the stage, and keys are deterministic — so resubmitting is running load --execute again.

The official SDKs

The pipeline depends on three first-party packages, and the split is worth knowing when reading the code:

  • @commercetools/importapi-sdk supplies every draft type in the plan. These are generated from the API's own specification, so a field that moves becomes a compile error rather than a rejected operation. Do not hand-write import resource types alongside them.
  • @commercetools/platform-sdk supplies the Project type and preflight's request builders.
  • @commercetools/ts-client owns transport: the client-credentials flow, token lifecycle, retry with backoff, a concurrency queue, and correlation IDs.

Three shapes are easy to get wrong by hand and impossible to get wrong with the generated types:

ShapeThe trap
PriceDraftImport.keyrequired — a price without one is rejected
Attributea discriminated union carrying type, e.g. {type: 'text', name, value}
a set attribute's type'text-set', not 'set'
The attribute discriminator is not derivable from the value: a string could be text, enum, lenum, date, datetime or time, and only the ProductType's declaration says which. That is why the mapper threads the ProductType's attribute definitions into variant mapping.
Two more fields are set deliberately rather than left to default: ProductDraftImport.priceMode (unset is assumed to mean Embedded) and sku, which the type marks optional — legal, but a migrated catalog needs one, so the gate reports a variant without it.
What the SDK does not own, and this pipeline still must: chunking into Import Requests of at most 20, container strategy, and deciding which Import Operation states to resubmit. That last one cannot live in transport at all — rejected and unresolved both arrive as HTTP 200 with the state in the body.

Credentials

From the environment or .env, never from the committed config. See .env.example.
VariableNotes
CTP_PROJECT_KEY, CTP_CLIENT_ID, CTP_CLIENT_SECRETrequired
CTP_REGIONderives all three hosts
CTP_AUTH_URL, CTP_API_URL, CTP_IMPORT_URLoverride individually
CTP_SCOPESoptional; omitting it grants every scope the API Client has
The Import API is on a different host from the HTTP API (https://import.{region}…). The conventional commercetools variable set has no import host, so it is derived from CTP_API_URL when absent.
Which project a run targets is refused rather than resolved, twice over. Exported variables otherwise win over the file, but not for project identity:
  • The file and the environment naming different projects is refused. The environment would win, so --env would silently not do what it looks like.
  • No file at all, while the environment names a project, is also refused. That is the case a disagreement check cannot catch — there is nothing to disagree with — and it is how a run in a directory with no credentials of its own borrows a project nobody chose for it. The only thing that catches it otherwise is preflight printing the project name, which is luck rather than a check.
    Set CTP_AMBIENT_OK=1 when the ambient environment is deliberate. CI and containers should; a working directory generally should not.
Preflight degrades rather than stopping. Without view_project_settings it cannot read the catalog model, locales, currencies or price countries — so it reports project-settings-unreadable (not project-unreachable, which means something else) and still runs what it can, which is the resource-count check that catches a load aimed at the wrong project. That needs only view_products.
Do not read a degraded pass as a pass. Everything preflight exists to verify depends on project settings, so a load in that state is going in blind on all of it — recoverable, because keys are deterministic and a re-run updates, but the recovery is manual. Add the scope and re-run. If you cannot, say so in the decision log with what was unverified, and prefer the dry run: it writes the exact request bodies to out/load-requests.json for review without sending anything.
Scopes: view_project_settings for preflight, manage_project_settings for --apply, manage_products plus manage_import_containers for the load. manage_products also grants the Import Requests for categories, product types, products, variants and embedded prices, and it covers reading and creating channels. Two more are conditional on the feed:
ScopeNeeded when
manage_standalone_pricespriceMode: standalone
view_customer_groups + manage_customer_groupsthe feed declares any customerGroup
view_product_selections + manage_product_selectionsthe feed declares any productSelection
view_stores + manage_storesthe feed declares any store
Neither is granted by manage_products. A missing customer-group scope fails the load at its first stage, before anything is imported — which is the cheap place to find out, but only if the scope list was not guessed.

verify

The last stage, and the only credentialed one that cannot change anything — every call is a GET.

npm run pipeline -- verify --config ../migration/migration.config.json
It answers the question load cannot: an accepted Import Request is not an imported resource. The Import API validates asynchronously, so a load can report every request accepted while operations are rejected afterwards — and because operations are retained 48 hours and counted per container, a container summary keeps reporting an old failure after a later run fixed it. Only reading the catalog settles it.
The 48-hour retention distorts the summary in both directions, and the second one is the dangerous half:
  • A stale failure that is already fixed still shows, because the old operation has not aged out yet. Reads as worse than reality.
  • A failure that has aged out is simply gone. Query a container more than 48 hours after the run and every surviving operation may well say imported — not because nothing failed, but because the ones that failed were deleted. Reads as better than reality, and it is indistinguishable from a clean run.
So a green summary is only evidence about a run you queried inside the window. Past that, the summary cannot answer the question at all and verify is the only thing that can, because it reads resources rather than operations and resources do not expire.
verify reads back exactly the keys the plan names, with a key in (...) predicate chunked at 100 keys per request, and reports:
  • missing — planned and absent. The load did not land, whatever it said.
  • mismatched — present but not what the plan says.
  • unexpected — a variant or attribute the plan never mentioned, which is how a stale earlier run becomes visible.

Two things it gets right that are easy to get wrong:

  • It compares staged data, not current. The pipeline imports unpublished on purpose, so current is empty on a first load.
  • It normalises read shapes first. An enum is written as a key and read back as {key, label}; references are written as keys and read back as ids. Enum values are normalised, and reference attributes are skipped rather than guessed at — a verifier that flags every enum in a catalog is one that gets switched off.
A failed read marks that kind unreadable and compares nothing, because a missing view_products scope and an empty project are otherwise the same observation.
Severity follows recoverability: a differing attributeConstraint is an error saying the attribute must be recreated (changeAttributeConstraint accepts only None), a differing price scope is an error saying the price must be deleted and remade (the Import API refuses to update country, customerGroup or channel), and a leftover variant is a warning.
--wait is not the same as resolved, and this is where that bites. load --wait drains processing; it does not wait out the 48-hour KeyReference window. So a load can return having reported every request accepted, with operations still unresolved — a category waiting for its parent, a product waiting for its category — and a verify run straight afterwards reports every one of those resources as missing, in wording written for a genuine failure.
Expect this at scale: a deep category tree can leave a hundred or more operations unresolved, and the resources they block absent, for minutes after the load returns — not hours, and not permanently. The deeper the tree, the longer the cascade.
verify now reads the operation states when it finds anything absent, and leads with operations-in-flight — the count still pending, and the instruction to wait and re-run before believing the absences. If the count does not fall between runs, something the plan referenced was never imported and the absences are real.

It reads stores and product selections back too, but only when the plan declares them — most engagements have neither, and an unconditional read costs two requests and two scopes for nothing. Two findings there are worth knowing in advance:

  • selection-mode-differs cannot be fixed by re-running. With no changeMode action, the selection has to be deleted and recreated, and every store referencing it re-wired. Until then the assortment means the opposite of the plan.
  • Store drift is reported per list, because the API replaces rather than merges each one. Supply-channel drift is only a warning: this pipeline imports no inventory, so that wiring affects nothing it loaded.
A store's key is matched verbatim, not prefixed — asking for mig-northwind-uk would find nothing and report the store missing.
Not built yet: orphan detection across the key prefix. verify checks what the plan names, not what the project holds that the plan has forgotten.

Configuration

migration.config.json. The dividing line for what belongs in it: if a different source or target project would need a different value, it is config; if it follows from how commercetools works, it is code.
Write it from the step-0 interview rather than copying a template — see SKILL.md. Five values are decisions, not settings: target.priceMode, productTypes.productLevelStrategy, productTypes.onMissingDefinitions, productTypes.searchableByDefault and keys.prefix. Each is irreversible or fails silently, so the loader refuses an absent or misspelled value with the consequence stated rather than falling through to one nobody chose — a generated config is not a trusted config.
A filled example ships at the package root as a reference and a fallback, not as the intended path. It carries keys.prefix: "REPLACE-ME", which the loader refuses, so a copy taken by mistake cannot run.

The field-by-field shape:

SectionKey fields
feeddir — directory of *.ndjson files
targetcatalogModel (Classic/Modular), priceMode (embedded/standalone; Modular allows only standalone)
marketdefaultLocale, requiredLocales, requiredCurrencies, currencyFractionDigits
keysprefix — every key becomes <prefix>-<sourceCode>
mediabaseUrl — required only when the feed has relative image URLs
productTypesonMissingDefinitions, defaultKey, defaultName, productLevelStrategy, searchableByDefault
loadbatchSize (max 20), maxOperationsPerContainer
The loader refuses, up front, four things that would otherwise fail deep into a run — or, in the first case, not fail at all: a Modular catalogModel, a currency with no currencyFractionDigits entry, a batchSize above the Import API's limit of 20, and a defaultLocale absent from requiredLocales.
Both catalog models are supported. What the loader refuses is Modular with embedded prices — Modular has no embedded prices at all, and a VariantImport has no price field, so that pair cannot be honoured whatever the pipeline does.
The catalog model decides the import shape and is fixed at map time: Classic emits a ProductDraftImport carrying its variants, Modular emits a container product plus separate VariantImport resources. Changing target.catalogModel therefore requires re-running plan — editing the config does not change a plan already written.

What shapes the load

Recorded here because it constrains the plan's shape, and verified against the Import API overview and best practices.
  • 20 resources per Import Request; fewer than 200,000 Import Operations per container; max 1000 containers per project (soft).
  • A container with no retentionPolicy is deleted 72 hours after creation. Set a TimeToLiveRetentionPolicy on create for a longer engagement.
  • Organise containers by resource type, not by temporal batch.
  • Import Operations are deleted 48 hours after creation, and an unresolved KeyReference resolves automatically if its target arrives inside that window — so prices may be imported before their products. Do not serialise the stages on polling; frequent polling of the summary endpoint actively slows the import. Retry only rejected; other states are retried internally, up to five times.
  • Import Resources accept only KeyReferences, never ids.
  • ProductVariantImport and StandalonePriceImport delete omitted fields on update. ProductVariantPatch is the only partial update.
  • Load in dependency order — channels and customer groups, then product types, then categories, then products, then Standalone Prices if the price mode calls for them — and tear down in reverse, scoped to keys.prefix. Channels and customer groups are outside that scope and outside the Import API entirely: created through the platform API, and removable only by hand.
  • Standalone Prices need manage_standalone_prices and customer groups need manage_customer_groups; manage_products covers every other request here.
  • Containers, Operations and Summaries are generally available; the per-resource Import Requests are public beta.

Fixtures

The pipeline's fixtures/ doubles as documentation of intended behaviour. Each one is runnable:
FixtureDemonstrates
declared-typesa clean catalog with a declared type system
inferred-typesthe same catalog with declarations stripped
derive-inferencedate, timestamp, set and multi-line inference
derive-conflictsthe three fatal ProductType clashes
derive-searchableisSearchable disagreeing across ProductTypes
derive-type-conflictone attribute name inferred as two types — refused, not warned
storesa store, a distribution and a supply channel, and an Individual selection with a SKU-scoped assignment
plan-edgesslug collisions and mixed-precision currencies
plan-money-erroran amount too precise for its currency
broken-schema, broken-integrityeach validation phase
audit-violationsa hand-written plan tripping every gate check
audit-truncateda half-written plan, to check the loader fails clearly
invalid-configthe four configuration combinations refused up front

Run any of them to see what a finding looks like before writing an adapter:

npm run pipeline -- validate --config fixtures/broken-integrity/migration.config.json
npm run pipeline -- audit --config fixtures/audit-violations/migration.config.json \
  --out fixtures/audit-violations/out
source-discovery.md

Source discovery

Profiling the real export — a sample of it, or all of it — before any mapping is designed, and turning it into a reviewable mapping proposal.

This is step 0b of SKILL.md. It comes before the remaining target decisions because it answers one of them and informs another: whether the source declares its own attribute types is a property of the export, not a question for the customer.

The rule

Profile the export, not the documentation about the source system. No two installations of the same platform agree on their own type system, and a migration designed against an assumed source model gets rewritten. The questions worth answering are in writing-an-adapter.md; this page is about answering them from data.

Do not teach the pipeline to read sources

There is an obvious-looking shortcut here that must not be taken: adding a discover command that parses CSV, XML, JDBC and whatever the customer exported. That is the universal source parser the whole architecture exists to avoid — see the source boundary in SKILL.md. Profiling is agent work. The pipeline's contribution is its existing inference, reached through a probe feed.

What to look for, and which decision it settles

Profiling produces a pile of observations unless it is organised by what the observations are for. Each row below is a thing to measure and the commercetools decision it decides:
Observe in the sampleSettles
One row per sellable SKU, or a parent/child hierarchy? How many levels?What is a product and what is a variant — the single largest modelling question. commercetools has exactly one level, so anything deeper collapses in the adapter
Which fields differ between rows sharing a parentCandidate variant axes
Which fields are constant across a parent's childrenCandidate SameForAll attributes
Whether each axis field has a stable code, or only display textWhether the adapter must derive codes. A localized value can never be variant identity
Distinct value count per fieldenum/lenum versus free text. A field with six distinct values across the whole sample is an enum; one with thousands is text
Whether labels accompany codesenum versus lenum — derive chooses on exactly this
Value kinds per field: all numeric, boolean-ish, ISO dates, mixedThe attribute type, and whether inference is even possible. Mixed kinds are a hard error, not a guess
Locale-bearing fields and their tag formatWhether the adapter must re-key en_GB to en-GB, and which locales actually appear versus which are configured somewhere
Price rows: currency, country, customer group, channel, validity, quantityPrice scope, and whether priceMode can even be honoured. Quantity breaks have no embedded-price equivalent
Whether a price is inherited from a parent levelcommercetools does not inherit between variants, so inheritance must resolve in the adapter
Decimal representation of money, and whether any amount has more places than its currency allowsWhether minor-unit conversion will refuse records
Category representation: path strings, parent references, adjacency listHow the adapter rebuilds the tree, and whether names repeat enough to force slug disambiguation
Maximum variants per parentThe catalog model. Classic caps at 100 per product — see 0a
Image paths: absolute or site-relativeWhether a CDN host has to be supplied from outside the export. Relative is fine in the feed — media.baseUrl resolves it, and validate refuses the pair without one
How many renditions per shot, and whether they are keyed by formatimages (one URL each) versus assets (one source per rendition). More than one rendition means assets, or the extras are discarded
Sentinel and placeholder values (N/A, -, 0000-00-00, empty strings that mean something)What the adapter must report rather than silently clean
Whether attributes live in more than one place in the exportMany platforms split attributes between a type definition and a separate classification structure. Mapping only the first half is a common and expensive mistake
Two of these deserve emphasis because getting them wrong is silent rather than loud: net versus gross prices is usually a setting elsewhere in the source rather than a property of a price row, and getting it backwards changes every price by the tax rate with nothing in the data to contradict you. And whether the product identifier is stable between exports — if it is not, keys derived from it duplicate instead of updating, and that has to be solved before anything else is designed.

When the handover is the whole export

Ask how large the export is before asking for a sample. If the whole thing is small enough to read — a few thousand lines, say — there is no sample, and the probe feed below is the wrong tool. Two reasons, and the second is the one that actually settles it:
  • Nothing to extrapolate to. The probe exists to predict what the full export will do; with the full export in hand, that prediction is just the answer.
  • If onMissingDefinitions is require, the pipeline's inference never runs at all, and inference is most of what the probe is for. A source that declares its own types — a hybris items.xml, a PIM schema export — takes that path, which is exactly the case where a complete small export is likely.
Do this instead: read all of it, then take the thin vertical slice from writing-an-adapter.md — one product, with its variants, prices and categories, all the way through audit. It catches the same class of problem the probe would, against real target invariants, and it becomes the first increment of the real adapter rather than a throwaway.

The rest of this page still applies. Profiling, the observation table, the mapping proposal and the sign-off are all independent of how much data you have; only the probe step is conditional.

Say which path you took in the mapping proposal. "Read the complete export, 8 products, 90 SKUs" and "profiled a 500-row sample of ~40,000" support very different amounts of confidence, and the reader cannot tell them apart afterwards.

The probe feed

For a large export, where a sample is all you will get up front. The part that makes discovery evidence rather than opinion.

Rather than reasoning about what types the source implies, write a throwaway adapter over the sample only, emit a probe feed, and run the offline stages on it:
cd ct-catalog-migration-pipeline
npm run pipeline -- validate --config probe.config.json
npm run pipeline -- derive   --config probe.config.json --out out/probe
Then read out/probe/MODEL-REVIEW.md.

What this buys that an agent's analysis does not:

  • The type, enum/lenum and attribute-constraint recommendations come from the same code that will run in production, so they cannot quietly disagree with it.
  • mixed-value-types fires on any field holding more than one kind of value — the case where inference is not merely uncertain but wrong for some records.
  • Every irreversible choice is already flagged: changeAttributeConstraint accepts only None, so a SameForAll or CombinationUnique proposed here is a one-way door, and derive marks it as such.
  • validate failures on the probe feed are themselves findings. An axis with no stable code, a localized value used as identity, a duplicated axis combination from a collapsed hierarchy — each is a real property of the source, surfaced before the real adapter exists.
  • It rehearses the adapter workflow at small scale, so the first real adapter is the second one written rather than the first.
Set productTypes.onMissingDefinitions to infer in the probe config even if the real run will use require: the point of the probe is to see what the pipeline would guess, and how confident it is entitled to be.
The probe adapter is disposable. It exists to produce evidence, not to become the real adapter — it handles only the sample, skips whatever is awkward, and should be deleted. Keeping it invites someone to grow it into the real one, which means the real one was designed against a sample.

The mapping proposal

Discovery ends in a named artefact, not a conversation, because it is the thing a human signs off. Write it next to the config:

# Mapping proposal — <engagement>

Sample: <what, how many records, from when>

## Source model as observed
<products vs variants, hierarchy depth, how categories are represented>

## Field mapping
| Source | commercetools | Type | Notes |
| :--- | :--- | :--- | :--- |
| ARTICLE_NO | variant sku | — | stable across exports; confirmed on two exports |
| COLOR_CODE / COLOR_NAME | attribute `colour`, axis | lenum | code + label present, so lenum |
| LONG_DESC_<locale> | product description | ltext | tags are `en_GB`; adapter re-keys to `en-GB` |

## Decisions needing sign-off
<each irreversible or lossy choice, with the cost of getting it wrong>

## Not migrating
<what is deliberately dropped, and why>

## Open questions
<what the sample could not answer>

The last two sections are the ones that get skipped and matter most. A migration claiming zero information loss is not being honest, and "not migrating" is where that honesty lives.

A sample is not the catalog

Skip this section if you read the complete export — the limits below are properties of sampling, not of discovery.

Otherwise everything above is a hypothesis. A sample under-reports exactly the properties that hurt:
  • Maximum variants per product — the constraint that can make a catalog unloadable is a long-tail property, and a 500-row sample will usually miss the worst case entirely.
  • Attribute cardinality — a field with six values in the sample and four thousand in production is text, not an enum, and the enum choice is irreversible.
  • Mixed value kinds — one bad record in fifty thousand is enough to make an inferred type wrong, and a sample is unlikely to contain it.
  • Sentinel values — rare by nature.
So discovery findings are confirmed against the full export when validate and derive run for real. validate recomputes the largest product's variant count and reports which catalog model it requires; derive re-runs inference over everything. Where the sample and the full export disagree, the full export wins, and the disagreement is itself worth reporting — it says something about how representative the sample was, which is worth knowing before cutover.

State this when presenting the proposal. A mapping proposal offered as conclusions rather than hypotheses is how a migration commits to an irreversible attribute constraint on the strength of 500 rows.

step-0-interview.md

The step-0 interview

Do not hand over a template to fill in. Ask the questions here, then write migration.config.json from the answers — the config is the record of a conversation, not a form.

Order of work

Size the export, then ask 0a, then discover properly — three steps, not two, because 0a needs a number only the source has:
  1. Size it (minutes): file list, line counts, largest variant count. Structural only — no mapping, no types, no decisions. Sizing is not the source work 0a means; writing an adapter is.
  2. Ask 0a with that number in hand. Both its answers can end the engagement, and both are cheaper to learn now than after an adapter exists.
  3. Discover the source (0b), then ask the rest (0c), because the data answers some of them.

0a. Can this pipeline serve the catalog at all?

Two questions, and they interact.

Both models are supported; the question is which this catalog needs. Ask it early — the answer fixes the import shape at map time, so changing it later means re-running plan:
AskWhy
Which catalog model is the project on, or intended to be — Classic or Modular?It decides which Import API resources are legal, and it is not inferable from data. A greenfield project can be switched either way with one setProductCatalogModel action, so it is a genuine project decision.
Roughly how many variants does the largest product have?Classic caps at 100 per product, Modular at 10,000. It is the one constraint that can make a catalog unloadable under Classic, and it costs one question to find out. 0b answers it from data if they do not know.

Resolve the two together, and say the outcome plainly:

ModelLargest productOutcome
Classic≤ 100Proceed.
Classic> 100The project has to move to Modular, or the catalog cannot be loaded. Say that before an adapter is written, not after.
ModularanyProceed. Over 100 is what Modular is for; at or under it Classic would also fit, so ask whether Modular is wanted for its own sake — it forces standalone pricing and has the defaultVariant caveat below.
An existing project already holding embedded variants is a different exercise: switching it is the staged Modular Catalog migration — dual-write, the InMigration state, a Variant copy job and a cleanup job — not one update action, and not this pipeline's job. Settle it before the catalog load rather than during it.
Two things to state whenever Modular is chosen, because neither is visible until late:
  • Pricing is Standalone only. Modular has no embedded prices at all, so target.priceMode must be standalone and the config refuses anything else. That also means the API Client needs manage_standalone_prices, which manage_products does not grant.
  • defaultVariant cannot be imported. Modular replaces the master variant with Product.defaultVariant, which appears nowhere in the Import API. The pipeline still chooses and logs it deterministically, then cannot write it: products load with no default variant, the plan records the loss, and setting them needs an HTTP API pass this pipeline does not do.
When a catalog cannot be served, say what is missing rather than implying the catalog is wrong. Do not suggest splitting products to fit Classic — that reshapes a catalog to suit a tool.

0b. Discover the source

Profile the export before the remaining questions, because the data answers some of them. Work through source-discovery.md — all of it, whichever path you take below.
Ask how large the export is first. If it is small enough to read whole — a few thousand lines — skip the probe feed, not the page: there is nothing to extrapolate to, and under onMissingDefinitions: require the inference it exercises never runs. Read all of it and take the thin vertical slice from writing-an-adapter.md instead.
Otherwise, for an export you only get a sample of: profile it, write a throwaway probe adapter over the sample, emit a probe feed, and run validate and derive on that. Type, enum/lenum and attribute-constraint recommendations then come from the code that will run in production rather than from an agent's reading, mixed-value-types fires where inference would be wrong rather than merely uncertain, and every irreversible choice arrives already flagged in MODEL-REVIEW.md.
Discovery ends in a mapping proposal — a named artefact, because it is what a human signs off. Its section-by-section template is in source-discovery.md; write it from there, not from this paragraph. It settles two things that would otherwise be guessed: the largest product's variant count for 0a, which a customer may genuinely not know, and whether the source declares its own attribute types — the onMissingDefinitions answer below.
Say which path you took, and if it was a sample, present the proposal as hypotheses, not conclusions — a sample under-reports exactly what hurts: worst-case variant counts, attribute cardinality, mixed value kinds, rare sentinels. "Read the complete export, 8 products, 90 SKUs" and "profiled 500 rows of ~40,000" support very different confidence, and a later reader cannot tell them apart.

0c. The remaining judgements

Now informed by 0b. Each is irreversible or fails silently, and a wrong-but-consistent choice passes every offline stage:

AskWritesWhy it cannot be defaulted
How does the storefront read prices?target.priceMode: embedded / standalonea product whose priceMode disagrees with where its prices are imports cleanly, reports imported, and shows no price at all
Which search API does the storefront use?productTypes.productLevelStrategy: sameForAll / nativenative Product-level attributes are invisible to Product Projection Search
Should attributes be searchable unless stated otherwise?productTypes.searchableByDefault: true / falseProductTypes that disagree on a shared attribute name make it unavailable for search, filters and facets everywhere — and the import still succeeds
Will this catalog be re-exported and re-loaded, or is this one-shot?nothing — DECISIONS.md onlyevery derived code becomes a key on the first load and cannot move afterwards. One-shot permits deriving from whatever is readable; repeatable requires deriving only from fields the source guarantees are stable, which is usually a smaller set
Which variant should be the shop window, if the source does not say?nothing — the adapter's isMasterabsent isMaster falls back to the lowest SKU by sort order. Deterministic, and arbitrary as merchandising: it decides what a category listing shows. Most sources have no master-variant concept, so this is the normal case, not an edge case
productTypes.onMissingDefinitions is not in that table: whether the source declares its own attribute types is a fact 0b established, not a preference. Set require when the adapter can emit attributeDefinition records, infer only when the source genuinely cannot describe its own type system — and then expect every inferred type to need review, because each one becomes an attribute constraint that cannot be changed afterwards.

The first two are storefront decisions with no evidence in the source, so they stay questions for whoever owns the implementation. The last two are the ones most often skipped — they feel like project management rather than modelling, and get discovered while writing the adapter, once the derivation is chosen.

0c-conditional. If slugs collide, ask who owns the URLs

Slugs are unique per locale across the Project, so a retail taxonomy that reuses names forces suffixes — and a forced suffix is a changed URL, which is an SEO decision and someone else's to make. derive and plan report the count; what they cannot do is find the owner. Ask while the mapping is still cheap to change, not after the load.
Two questions, not one: the policy, and the name. Asking only the policy logs every changed URL against nobody, which reads as a decision and is not one. The name is what makes the entry a decision rather than a guess.

Do not raise it when nothing collides.

0c-conditional. If 0b found relative media URLs, ask about the host

Do not raise this otherwise — it is not a standing question.

Sources commonly store /medias/... and keep the host in a CDN setting, a storefront config, or an operations runbook. It is not in the export, so it can only be supplied. Put it to the user plainly, with the cost: a wrong base loads a catalog whose every image 404s, and neither the pipeline nor the project can detect that — the only way to find out is to open one.
Ask what host or prefix these image paths resolve against; it writes media.baseUrl. Two acceptable answers. A confirmed prefix — set media.baseUrl, and the mapper resolves every relative URL against it, recording the count and the base as a reviewable decision. No host available — drop the images in the adapter and report them as not migrated. An absent image is recoverable; a wrong URL on every product is not, because it looks like success. Do not guess: validate refuses relative URLs with no media.baseUrl for exactly this reason. Absolute URLs need none of it, and a feed may mix the two.

0d. The project facts

Lookups rather than judgements: keys.prefix (name this engagement — it goes into every key created, and bounds a teardown to its own work), market.defaultLocale, requiredLocales, requiredCurrencies, and a currencyFractionDigits entry per currency. preflight checks the market values against the real project later; approximate is fine now.
Then write the config. The field-by-field shape is in running-the-pipeline.md; the pipeline repository's root migration.config.json is a filled example, a fallback rather than the intended path — it ships keys.prefix: "REPLACE-ME", which the loader refuses, so a copy taken by mistake cannot run.
Much of this is intent, or a hypothesis from a sample — not fact, and the later stages confirm all of it: validate recomputes the largest product's variant count from the whole feed and reports which model it actually requires, derive re-runs inference over every record rather than 0b's sample, preflight reads the project's real productCatalogModel. Where intent and evidence disagree the evidence wins, and the disagreement is the finding — it says how representative the sample was, before cutover instead of after. The loader still refuses an absent or misspelled value with the consequence spelled out: a generated config is not a trusted config.

Then open the decision log

The config records what was chosen, not who chose it, what the alternative would have cost, or that a question was asked at all. Write one entry per decision from 0a and 0c, plus any question that could not be answered — an open question with an owner is a plan, an unasked one is a surprise. Format and placement: decision-log.md.
Keep writing it at every step from here, not at the end: a log reconstructed afterwards records what someone remembers deciding, reliably the subset that turned out well, while the entries that matter are the irreversible ones taken under uncertainty.
writing-an-adapter.md

Writing a source adapter

The adapter is the only code an engagement normally writes. It reads the source and emits NDJSON against the feed contract. It is deliberately small and deliberately disposable — typically a couple of hundred lines.

Everything source-specific belongs here. Nothing source-specific belongs downstream.

Establish what the source is before writing any mapping

A migration designed against an assumed source model gets rewritten. Answer as much of this as the export allows, from the export itself rather than from documentation about the source system.

source-discovery.md is the method for answering these from data rather than by hand — profile a sample, emit a throwaway probe feed, and let derive propose the types. Work through that first; this list is what it is trying to answer.
Coverage
  • How many products, variants, categories, and prices? Orders of magnitude, not exact counts — it decides whether the load design matters.
  • Is this the whole catalog or a sample? A sample that behaves differently from production is worse than no sample.
  • Which locales and currencies actually appear in the data, as opposed to which are configured somewhere?
Identity
  • What is the stable machine identifier for a product? For a sellable SKU?
  • Can it change between exports? If yes, keys derived from it will duplicate rather than update, and that has to be solved before anything else.
  • Is there a separate identifier downstream systems join on? It goes in externalId.
The type system
  • Does the source declare attribute types, or are they implied by values? Declared means emitting attributeDefinition records; implied means the pipeline has to infer, and every inference needs review.
  • Are attributes split across more than one place? Many platforms keep some attributes in a product type definition and most of them in a separate classification or taxonomy structure. Mapping only the first half is a common and expensive mistake — check for a second source before designing anything.
  • Which attributes vary per variant, and which are invariant per product?
  • Which attributes constitute variant identity?
Variants
  • How many levels does the hierarchy have? commercetools has exactly one, so anything deeper collapses in the adapter.
  • What distinguishes two variants of the same product? Is that value a stable code, or display text?
  • Are there non-sellable intermediate levels? They usually become nothing.
Prices
  • Is a price inherited from anywhere? commercetools does not inherit between variants, so inheritance resolves here.
  • What dimensions scope a price — currency, country, customer group, channel, date range, quantity?
  • Are prices net or gross? This is often a setting elsewhere in the source, not a property of the price row, and getting it backwards changes every price by the tax rate with nothing in the data to contradict you.
  • Quantity breaks have no embedded-price equivalent. Expect to report them.
Media
  • Are image paths absolute or site-relative? If relative, the host is not in the export — it is a question for whoever owns the storefront or CDN.
  • How many renditions per shot? More than one means assets, not images.
  • Does the source carry real dimensions, or will they default to 0×0?
Categories
  • Is the tree navigation, or is it really a set of facets? A taxonomy used for filtering is usually better as attributes.
  • Are names unique? Retail taxonomies reuse them across departments, which forces slug disambiguation.
  • Is there a sibling ordering worth preserving?

Shape of an adapter

read source  →  assemble  →  emit
   parse         rebuild      one NDJSON
   the grammar   hierarchies  record per line
                 resolve
                 inheritance
Split parsing into syntactic (the file grammar — headers, cells, encodings) and semantic (what the rows mean — rebuilding hierarchies, attaching prices at the level they were declared). Conflating the two makes both untestable.

Emit in any order. The contract allows forward references, so there is no need to buffer the whole catalog to get ordering right.

Rules that matter

Drive type coercion from the source's declared types, never from a list of field names. A name allowlist — "these four fields are dates" — fails silently on the next export: the value lands as the wrong JSON type and commercetools rejects the whole record with a terse error.
Emit attributeDefinition records if the source declares types at all. It is the difference between a translation with a defensible origin for every choice and a pile of guesses needing sign-off.
Never repair source data silently. If a cell holds a sentinel value, a placeholder, or obvious junk, either skip it and report it or pass it through — quietly cleaning it is how a migration starts lying about what the source contained.
Report every silent cap. If the adapter samples, truncates, or skips, say so in its output. Partial coverage that reads as complete is worse than an error.
An attribute the source declares but never populates: emit the definition, not the attribute. Declaring it keeps the target's model faithful to the source's, which matters if the catalog is ever re-exported into it; omitting it avoids a column no product fills. The pipeline reports either way — attribute-never-populated if you declare it — so both are defensible and the cost of choosing wrongly is low. Pick the first by default, and say which you did, because a later reader comparing the two models will otherwise wonder whether the field was missed.
Normalise locale tags. Many sources write en_GB; commercetools requires en-GB. Re-key every localized map, and make sure the default locale is populated — a localized field with no entry for the default locale renders empty.
Verify the encoding from the bytes; a declaration can be wrong. An XML prolog saying latin1, or an export doc claiming a codepage, is a claim about the file, not a fact. Trusting encoding="latin1" on a file that is actually UTF-8 puts °C into the attribute label — that exact signature is what to look for, and MODEL-REVIEW.md is usually where it first becomes visible. file * costs nothing, and mis-decoded text survives every validation in this pipeline — it is well-formed, correctly typed, and wrong. Check the degree signs, the accented names and the currency symbols in the first artefact you generate.
Derive a code when the source only has a display value. Especially for variant axes. Record how it was derived; that is a decision, not a detail.
Choose the master variant deliberately, because the fallback is arbitrary. With no isMaster, the pipeline takes the lowest SKU by sort order and records the choice for review. That is deterministic on purpose — a feed-order-dependent master would silently change between runs — but it is not a merchandising decision, and the master variant is what a storefront shows by default and what a category listing puts in the shop window.
This is not an edge case. Most source systems have no master-variant concept at all: in a hybris export the base product is the style, so every product gets its master by sort order — which is how a dress ends up fronted by Navy XXL. The dry run's load-requests.json is where that is visible before the load, and usually the only place.
If the source has any signal — a display order, a hero image, a "default colour" — map it to isMaster. If it has none, that is a question for whoever owns merchandising, and it is cheap to ask at step 0 and expensive to change after the first load.

Media

Images are a normal migration concern and they have three decisions in them, none of which the data answers.

Emit the URL as the source stores it, relative or absolute. Do not resolve it in the adapter and do not invent a host. Sources very often keep a site-relative path — /medias/sys_master/root/h9c/... is the hybris shape — and the host lives in a CDN setting or a storefront config, not in the export.
The pipeline owns this. validate refuses a feed carrying relative URLs unless media.baseUrl is set, so the decision is forced into the open instead of being guessed; plan then resolves each relative URL against that base and records the count and the base as a reviewable decision. Absolute URLs pass through untouched, and a feed may legitimately mix the two.
Do not work around this by inventing a hostname to get a feed to validate. That is what a contract requiring an absolute url would force, and it is precisely the guess the current rule exists to prevent.

If no host can be obtained, drop the images in the adapter and report them as not migrated. An absent image is recoverable; a wrong URL on every product is not, because it looks like success.

Several renditions per shot go in an asset, not an image. Sources commonly hold thumbnail, product and zoom behind a container or a format table. An image carries exactly one URL, so using it means discarding the rest; an asset holds one source per rendition, each with its own key, dimensions and content type. A hybris MediaContainer maps to one asset, and nothing is lost.
Use images when the source genuinely has one URL per shot. Use assets the moment it has more, and key the sources after the source's own format names (thumbnail, product, zoom) rather than inventing a scheme — the keys are how a storefront asks for a size.

If you do have to reduce several renditions to one image, larger is the safer default: a storefront can scale down and cannot scale up. Record the choice — "the middle one" is a decision a retina storefront will disagree with.

Dimensions are required, and 0×0 is accepted. If the source has real width and height, emit them. If it does not, the pipeline defaults to 0×0 and records it as information loss — a storefront reserving layout space from the declared size then cannot. That is worth knowing before go-live rather than after a page reflows.
Two smaller things worth stating in the proposal: image order is significant in commercetools and usually meaningful in the source, so preserve it rather than emitting whatever the join returned; and label is optional but is what alt text comes from, so dropping it is an accessibility cost, not a cosmetic one.

Iterating

Run validate after every adapter change. Its diagnostics name a file and line in the feed, and they are written to be actionable:
ERROR catalog.ndjson:17 [axis-combination-duplicate]
      Variants 'COLLAPSED-M-1' and 'COLLAPSED-M-2' share the axis combination
      size=M. Axes become CombinationUnique, so commercetools will reject the
      product. Either the hierarchy was collapsed wrongly, or an axis is missing
      from the declaration.

Fix the adapter, never the feed. The feed is regenerated on every run.

A useful first milestone is a thin vertical slice: one product with two variants, two prices and a category, all the way through audit. It exercises every stage and every invariant in seconds, and it surfaces identity and type-system problems while they are still cheap.

What good output looks like

  • Codes are stable machine identifiers, and externalId preserves the raw source identifier.
  • Axis values are codes; display text is in axisLabels.
  • Every SKU has an explicit price, with inheritance already resolved.
  • Locale tags are IETF form and the default locale is populated everywhere.
  • Money is a decimal string with no more precision than its currency allows.
  • Anything the adapter could not express is reported rather than approximated.