Deep dive · Hector AI and catalogue content

Hector AI and Initpc's product pages: clean data, controlled text

Thin, inconsistent and duplicate supplier product pages on a catalogue nobody could rewrite by hand. Hector AI normalised the data before the text and generated titles and descriptions with checks on every batch.

  • 4.7 · 38 Google reviews
  • 25 years of experience
  • ISO 9001 / ISO 27001
Initpc.it category page with its descriptive text and first products: catalogue content generated and checked with Hector AI
160,000 product pages generated and checked

01 / The challenge

A huge catalogue, described by supplier price lists

Around 160,000 products imported from several suppliers and warehouses, each with its own way of writing titles and attributes. On the site the result was thin product pages, inconsistent across products of the same family and identical to those of dozens of other shops. Writing them by hand was impossible, and the estimate with commercial AI APIs exceeded €30,000, to be paid again with every regeneration.

  • Pages copied from the price list: nothing on compatibility or use
  • Attributes scattered across titles, abbreviations and fields that vary by supplier
  • Duplicate content Google had no reason to index

02 / The solution

An in-house pipeline on Hector AI: data first, then text, then checks

Hector AI normalises the raw data, distils the attributes into structured fields and only then generates titles, descriptions, HTML structure and keywording for each product page. Every batch goes through validation against the ERP, duplicate checks, blocking rules and human sampling before publication.

Tools and technologies

Technologies and tools we used

  • Hector AI on our own infrastructure
  • Product data normalisation
  • Attribute distillation into structured fields
  • Generation of titles, descriptions and HTML
  • Semantic keywording per product page
  • Validation against the ERP
  • Batch publishing with human sampling

03 / The results

What we achieved

  • 160,000 product pages generated and checked with Hector AI in the first campaign
  • 30,000 euros estimated with commercial APIs, avoided by building the pipeline in-house
  • 280,000+ products in the catalogue today, with data normalised at the source
  • Within a few weeks, a significant increase in organic visibility and traffic on the rewritten categories
  • Long-tail categories and brands indexed for precise searches and compatibility queries
  • Structured attributes reusable for filters, marketplace feeds and the future proprietary search

04 / Technology partners

  • 25+ Years supporting businesses
  • 4.7/5 Average rating on Google
  • 38 Verified reviews
  • ISO 9001/27001 ISO certifications

The product page is the most numerous page on an e-commerce site, and the most neglected. On a catalogue like Initpc's — stationery, office supplies, IT, school and gifts, now over 280,000 SKUs — it is where the customer decides whether to buy and where Google decides whether that page deserves a place in the index. When product pages are copies of the supplier's price list, both decisions go badly.

Here we describe how we used Hector AI, our artificial intelligence platform, to rewrite the catalogue starting from the data rather than the text: attribute normalisation, controlled generation of titles and descriptions, checks before publication, and close attention to the risk of mass generation — thousands of thin pages that a search engine learns to ignore.

It is one piece of the full architecture of the Initpc project: the data comes from the ERP, and the distilled attributes are destined to become fields in the search engine.

The starting point: thin, inconsistent and duplicate product pages

When we tackled the problem the catalogue held around 160,000 products, imported from several suppliers and warehouses. Every supplier describes its items in its own way: the colour in the title or in a separate field, the size as an abbreviation, the pack quantity to be guessed from the price. The same pen, arriving from several price lists, comes with different titles and descriptions that say the same things in a different order.

On the site the result was predictable: thin product pages (the price-list title, a single line, nothing on compatibility or use), inconsistent product pages (same product family, different structures) and duplicate product pages, identical to the supplier's and therefore to dozens of other shops.

Writing them by hand was not an option, and neither was rewriting "just" the most important ones: the value of a catalogue like this lies in the long tail, in the items people search for by code or by a precise characteristic. There, a well-written product page brings a qualified visit, and a copied one brings nothing.

Buying the text or building the pipeline

The first route we evaluated was the obvious one: generating the product pages through the APIs of commercial AI services. We ran the numbers on the real catalogue, step by step, and the estimate came to more than €30,000 in API spend alone — to be paid again with every regeneration.

We chose the opposite route: building the pipeline in-house on Hector AI, which runs on our own infrastructure with open source models under our control. The cost becomes that of energy and servers, which we already manage; the catalogue — prices, suppliers, margins — never leaves the company's perimeter; and a batch can be rerun as many times as needed. How Hector AI works is covered on its own page: here we are interested in the method.

Data first, then words

The most common mistake in AI projects for product descriptions is asking the model for a description starting from a messy title. The model writes, confidently, and fills the gaps with whatever seems plausible. That is why the first phase generates no text at all: it normalises.

Normalisation

Hector AI reads titles, descriptions and raw fields and brings them into a uniform form: consistent units of measure, abbreviations spelled out, clean capitalisation and punctuation, codes aligned with the ERP. It works alongside what the ERP does on supplier price lists for prices and stock, with a different goal: there the data is normalised to sell, here to describe.

Attribute distillation

From the normalised text Hector AI extracts the attributes that matter for that category and writes them into structured fields: brand, size, colour, pack type and quantity, material, compatibility (the printer for a cartridge, the model for a case), intended use. Each attribute keeps its source, so a later check can trace it back to the point in the price list that justifies it.

This is the project's real asset: the text can be regenerated, while clean attributes serve product pages, filters, marketplace feeds and search.

Generation and checks

Only at this point does generation come in, with a rich context: distilled attributes, category, similar products, editorial rules. For each product page Hector AI produces a title following a per-category pattern (brand, type, distinguishing feature, size or pack); a description that starts from the verified attributes and ties them to real-world use, adding nothing that does not come from the data; a consistent HTML structure — introduction, feature list, notes on compatibility and packaging — shared across the whole product family; and semantic keywording, that is, the expressions people actually use to search for that product, used to link the page to its category and brand.

Descriptions have to be accurate, complete and different from the supplier's, and no product page goes live without these checks:

  • Attribute validation against the ERP. Every attribute mentioned in the text is compared with the ERP data: if the description indicates a pack size different from the known one, the page is held back.
  • Duplicate check. Titles and descriptions are compared with each other and with the suppliers' text; pages that are too similar are regenerated or merged.
  • Blocking rules. Forbidden expressions (superlatives, promises, prices or stock levels that change), minimum and maximum lengths per field, mandatory HTML structure.
  • Human sampling. From every batch a share of the product pages is read by a person, above all in technical categories where a compatibility error costs a return; if a systematic flaw emerges, the whole batch goes back.

The risk of mass generation and how we contained it

Generating tens of thousands of pages in a short time is exactly what search engines have learned to penalise, and rightly so: most "AI-enriched" catalogues are made of interchangeable text. We treated the risk as a project requirement, not a side effect.

Mass generation risk Control applied
Thin content: long but empty descriptions, identical across different products Generation from distilled attributes; a minimum threshold of information, not of characters
Content duplicated from suppliers and other shops Comparison with the original price-list text; regeneration of pages that are too close
Hallucinated technical characteristics (compatibility, sizes, quantities) Validation of every attribute against the ERP; page blocked if a data point cannot be confirmed
Penalties for sudden mass publication Publication in batches by category, watching indexing before the next batch
Effort spread across low-value categories Priority to the categories with the highest margin, the most searches and the longest tail

The batches deserve a note: we started from the categories with the most commercial value, watched how Google re-crawled and re-indexed them, and only then moved on to the next ones, correcting the rules on the first product families before applying them to the others.

What changed

The signal came quickly: within a few weeks of the first batch we observed a significant increase in organic visibility and traffic on the rewritten categories. We do not publish percentages: a number without its seasonal and commercial context says very little.

Long-tail category and brand pages — previously crawled but not indexed — began to appear for precise searches, by product code or technical characteristic. Product pages with complete attributes latched onto compatibility queries, which for cartridges, toner and accessories are the main way people search, and consistent titles made results readable across several variants of the same item.

Where we are heading: search and commercial Hector AI

The distilled attributes do not stop at product pages. Initpc's search currently relies on Elasticsearch through ElasticPress; the new proprietary engine we are developing will use brand, size, colour, pack and compatibility as index fields and ranking criteria, so that someone looking for cartridges for a given model finds the compatible products rather than those that merely mention it. This integration is in development, not in production.

In the same workshop is commercial Hector AI with a human operator in the loop: an assistant that helps the customer choose between similar products, suggests alternatives when an item is unavailable and hands the conversation over to a person when the request goes beyond the catalogue. This too is in development, not in production.

It is the path we propose in our AI solutions for retail and e-commerce: data first, then text, then services built on the data. It is the part of the Initpc case study that transfers most readily to other catalogues.

Frequently asked questions

The questions we get about this project

Can AI write an e-commerce site's product descriptions without inventing characteristics?

Yes, provided it starts from the data rather than the text. Every attribute that appears in a product page was first extracted into a structured field and verified against the ERP; generation works only on those fields, and a page that cites a data point with no confirmation is blocked before publication. Hallucination is not eliminated by asking the model to be careful, but by removing its ability to invent.

Doesn't generating thousands of product pages with AI risk a Google penalty?

The risk exists when the pages are thin or duplicate content published all at once. We reduced it by working in batches by category, checking similarity between pages and against the suppliers' text, sampling every batch with a human read-through and giving priority to the categories with the most value. Google does not penalise AI as such, but pages that add no information: with accurate, complete descriptions that differ from the supplier's, the effect we observed was the opposite.

Why build your own pipeline instead of using the APIs of commercial AI services?

For the Initpc catalogue the estimate with commercial APIs exceeded €30,000, to be paid again with every regeneration. With Hector AI, which runs on our own servers with open source models, the cost comes down to energy and infrastructure, batches can be rerun as many times as needed and commercial data — prices, suppliers, margins — never leaves the company. For a catalogue that changes constantly, that is worth more than the initial saving.

More case studies

Thousands of product pages to write and nobody to write them?

Poor supplier data, duplicated descriptions, categories without text: tell us about your catalogue and we will tell you what Hector AI can do — and what we would not let it do.

06 / Talk to an expert

Talk to the people who built the Initpc architecture

Tell us about your catalogue, your platform and what no longer holds up: we call you back within one business day.

A senior consultant reviews the situation with you — infrastructure, search, data management, migration — and proposes the solution proportionate to the real scale of your store, with no obligation.

At least 10 characters.

Fill in to send: Name, Email, Message, privacy consent.

I have read the Privacy Policy and consent to the processing of my data.

We use cookies

We use technical cookies required for the site to work and, only with your consent, analytics and marketing cookies. You can accept, refuse or choose category by category. If you continue browsing to another page without choosing, cookies are considered accepted. Cookie Policy