AI and customs: how classification
technology is evolving.
Product classification has gone from printed tariff books to keyword search tools to systems that can read a product description and reason through the same rules a trade specialist would apply by hand. Here's what actually changed, what today's AI-driven tools are genuinely good at, and where a person still needs to be in the loop.
8 min read · Trade technology
From tariff books to language models
Product classification has always been a research task dressed up as a paperwork requirement. For most of its history, doing that research meant a person with a printed or digital tariff schedule, working through headings and notes by hand, product by product. The tools around that process changed long before the process itself did.
The first wave of software support was essentially search: digitized tariff schedules with keyword lookup, letting a reviewer find candidate headings faster than flipping through a printed volume. That helped with speed but did nothing for judgment. The reviewer still had to read the notes, apply the rules, and decide which heading actually fit.
A second wave added rules engines and decision trees: if a product has this material and that function, route it to this heading. These systems could automate genuinely simple, well-defined cases, but they were brittle. Any product that did not fit neatly into a pre-built branch fell back to manual review, which meant the hardest products, the ones most likely to be misclassified, still got the least automated support.
The current wave, built on large language models, is different in kind rather than degree. These systems can read an unstructured product description, the kind a person might actually write, and reason through the same chain of headings, notes, and interpretive rules a trained classifier would apply, rather than requiring the input to already be sorted into pre-built categories.
What actually changed, concretely
The shift from rules engines to language models sounds abstract, but it shows up in a few very practical ways.
| Capability | Rules engines | Language model-based systems |
|---|---|---|
| Input format | Requires structured, pre-categorized fields | Can work from a plain-language product description |
| Handling novel products | Falls back to manual review | Can reason through unfamiliar products against the same rules |
| Explaining a result | Traces a fixed decision path | Can generate a written rationale citing the specific rule applied |
| Adapting to schedule updates | Requires manually rebuilding decision branches | Can be updated by supplying the revised schedule text directly |
| Multi-country schedules | Needs a separate rule set per country | Can apply a shared reasoning process across schedules |
None of this means the newer systems are simply "smarter" in some vague sense. It means the bottleneck moved. The old bottleneck was building enough branching logic to cover every product variation in advance. The current bottleneck is closer to what it has always been for a human classifier: having an accurate, specific enough product description to reason from in the first place.
Where AI genuinely helps today
Working from real product descriptions
A system that can read "a padded nylon bag with a rigid internal frame, designed to carry laptops and books" and reason from that, rather than requiring the input already tagged as "backpack, structured, laptop-compatible," removes a whole layer of manual pre-processing.
Producing a documented rationale automatically
A written explanation of which rule was applied and why, generated at the same time as the code itself, solves one of the most common gaps in manual classification: reasoning that exists only in a reviewer's memory.
Applying the same logic across many destination schedules
Because the underlying reasoning process is not hard-coded per country, the same system can be pointed at different national schedules and apply a consistent method, rather than needing a separately built rule set for each one.
Surfacing its own uncertainty
A well-built system can flag when a product is genuinely ambiguous or sits close to a boundary between two headings, rather than confidently returning a single answer regardless of how contested the case actually is. That flag is what makes routing to human review possible at volume.
Where it still falls short
It is worth being direct about the limits, since overstating what these systems can do is its own kind of risk in a compliance context.
- A model can only reason as well as the product description it is given. Vague or misleading input still produces an unreliable answer, no matter how capable the underlying reasoning is.
- Genuinely novel or legally contested products, the kind that sometimes require a formal binding ruling from a customs authority, still need expert human judgment and are not the right use case for automation alone.
- Model outputs can sound confident even when they are wrong, which makes a confidence signal and a review process more necessary, not less.
- Tariff schedule updates still need to be sourced and fed to the system accurately. Automation does not remove the need to track when a schedule has actually changed.
The role of human review, going forward
None of this points toward removing people from the process. It points toward changing what people spend their time on. In a manual process, a reviewer's time goes toward every product equally, regardless of how routine or contested it is. In an AI-assisted process, that time can be concentrated on the smaller set of products that are actually ambiguous, novel, or high-stakes, while routine, high-confidence matches move through with a lighter check.
This is less a replacement of expertise and more a redistribution of it. The trade specialist who used to spend most of a day on routine classifications can spend that same day on the handful of genuinely hard calls, with the system handling the volume around them and surfacing the cases that need attention.
How customs authorities are responding
Customs authorities have generally not objected to automated classification tools; what they have consistently asked for is documentation. A determination needs to be explainable and defensible after the fact, regardless of whether a person or a system produced it. This is one reason the shift toward systems that generate a written rationale alongside the code matters beyond convenience: it is closer to what an audit actually requires than a code with no supporting reasoning attached.
The practical expectation, in most jurisdictions, remains the same as it has always been for classification generally: the importer or exporter is responsible for the accuracy of the code, regardless of what tool produced it. Automation changes how the work gets done. It does not change who is accountable for the result.
Looking ahead
The trajectory so far suggests the technology will keep moving from narrow, rules-based automation toward systems that can reason more broadly across product data and schedule text, with a shrinking but still essential role for human review on the hardest cases. What is unlikely to change is the underlying standard: a defensible code, with documented reasoning, that holds up if it is ever questioned. The tools for getting there keep improving. The bar for what counts as a correct answer has not moved.
See where AI-driven
classification actually stands.
Enthron applies AI reasoning to classification, complete with a documented rationale for every code and a confidence signal that routes ambiguous products to human review.