Blog · Getting started

How AI improves
HS classification accuracy

Manual classification works fine at small volumes, but it starts to strain the moment a catalog grows past a few hundred SKUs or ships into more than a handful of countries. This guide covers where classification errors actually come from, how an AI system approaches the same problem, and where it genuinely outperforms manual review versus where a person still needs to be in the loop.

8 min read · Trade compliance basics

Where classification errors actually come from

Manual classification depends almost entirely on the person doing it: how much product knowledge they have, how carefully they read the heading notes, and how much time they have for any single item. At small volumes that works fine, because someone can spend real time on each product.

The strain shows up as volume grows. A catalog of a few hundred SKUs is manageable for a trained team. A catalog of tens of thousands, updated weekly, shipped into a dozen countries each with its own tariff schedule past the sixth digit, is a different problem entirely. Reviewers get faster by necessity, and speed is usually where consistency slips.

Most classification errors are not exotic edge cases. They come from skipped steps: classifying by product name instead of composition and function, missing a chapter note that changes the answer, or applying yesterday's logic to a product whose specs quietly changed. None of these require an unusual product. They just require enough volume that something eventually gets rushed.

What actually changes when AI does the work

An AI classification system does not get tired, does not skip steps under deadline pressure, and applies the same reasoning to the thousandth product of the day as it did to the first. That alone addresses the most common source of manual error, but it is not the only difference.

DimensionWhat changes
ConsistencyThe same reasoning path applies to every product, regardless of who "would have" classified it or when.
Speed at scaleA catalog of thousands of SKUs can be classified in the time a person would spend on a handful.
Rule trackingSchedule updates and chapter note changes can be reflected the moment they take effect, not whenever someone gets around to reading the bulletin.
Audit trailThe reasoning behind each code is logged automatically, rather than depending on someone remembering to write it down.
Judgment on genuine edge casesStill generally needs a person: essential character calls, novel products, and disputed classifications benefit from human review.

That last row matters. AI classification is not a replacement for judgment on the hardest cases, it is a way to stop burning that judgment on the routine ninety percent of products where the answer is not actually in question.

Anatomy of an AI classification pipeline

A well-built classification system does not just guess a code from a product title. It works through a structured pipeline that mirrors the same legal reasoning a trained classifier would use, just applied consistently and at speed.

CLASSIFICATION PIPELINE Product data INTAKE Feature EXTRACTION Candidate HEADINGS GRI rule APPLICATION Confidence SCORE Name, spec, material Function, composition All plausible headings Rules 1–6, in order Flags weak matches FULLY AUTOMATED ROUTES TO REVIEW

The confidence score at the end is the piece that keeps the system honest. A product with a clear, well-documented match moves straight through. A product where the system's own reasoning is thin, contradictory, or split between two plausible headings gets flagged for a person to look at, rather than silently guessing.

How it works end to end

In practice, running a product through an AI classification system looks like a short, repeatable sequence, whether it is one SKU or an entire catalog upload.

01 Product data is submitted: description, spec sheet, or an existing catalog 02 The system identifies material, function, and construction from the data 03 It applies the General Rules of Interpretation to reach a candidate code 04 Low-confidence results are routed to a person for review, not auto-accepted 05 Final code and reasoning are logged against the product for future audits

Step four is the part that separates a genuinely useful system from a black box. The goal is not to remove human judgment entirely, it is to spend that judgment only where it is actually needed, on the small fraction of products where the classification is genuinely unclear.

Common mistakes when adopting AI classification

01

Treating it as a black box and skipping the reasoning

A classification tool that returns a code with no explanation is not actually more defensible than a guess. Look for a system that shows which rule it applied and why, not just the final number.

02

Auto-accepting low-confidence results

The value of a confidence score comes from acting on it. Routing weak matches to a person is the whole point; treating every output as final defeats the safeguard the system was built with.

03

Feeding it thin product data

An AI system can only reason from what it is given. A product name with no material, function, or spec detail produces the same weak result an underinformed person would produce, just faster.

04

Assuming one classification covers every destination

A good system classifies per country, since the first six digits are shared but the rest of the code, and the applicable duty rate, are not. A tool that returns a single global code for every market is skipping a real step.

Where it helps most

AI classification is not a universal replacement for manual review, it is a better fit for some situations than others. Matching the tool to the task is what actually improves accuracy.

SituationBest fit
Large or fast-growing catalogAI, with routine review of flagged items
Shipping into many countriesAI, since it can apply country-specific schedules consistently
Novel or genuinely ambiguous productPerson, ideally supported by the AI's candidate headings
Product tied to a disputed or prior rulingPerson, with the ruling as the controlling reference
Routine catalog maintenance and updatesAI, to keep pace with schedule changes automatically

Most mature setups end up as a hybrid: AI handles the bulk of routine classification and flags anything genuinely uncertain, while a person reviews only the flagged items. That division tends to produce both faster throughput and better accuracy than either approach running alone.

Getting started

If your company is evaluating AI classification for the first time, a few things are worth checking before committing to a tool:

  • Ask to see the reasoning behind a sample classification, not just the resulting code.
  • Confirm it classifies per destination country, not just to a shared six-digit base.
  • Check how it handles low-confidence results, and whether those genuinely route to a person.
  • Feed it a few of your hardest, most ambiguous products first, not just the easy ones, to see how it actually performs.

The goal is not to remove judgment from classification, it is to stop spending that judgment on the products that never needed it in the first place, so it is available for the ones that do.

Next step

See AI classification
applied to your own catalog.

Enthron classifies against national tariff schedules across dozens of countries, flags low-confidence results for review, and logs the reasoning behind every code.