️ Documentation · Collect

Setting Up Collect

How to point Trinis at your vendor's catalog, understand automatic platform detection, parallel collection, and the cache that avoids reprocessing what hasn't changed.

Last updated: September 2, 2026

Collect is the first stage of the pipeline (Collect → Enrich → Deliver): it reads the vendor's catalog and delivers a list of structured products for enrichment. All the configuration lives in a vendor — you set it up once and reuse it for every job.

How Collect works

When a job starts, Collect:

  1. Uses the listing URL configured on the vendor.
  2. Walks through the listing pages in sequence, respecting pagination, until it reaches the job's product limit.
  3. Fetches each product's detail pages in parallel and extracts EAN, name, price, description, spec sheet, and images.
  4. Discards products without a valid EAN (12 to 14 digits) — the EAN is the key that links the product to the store catalog and to the cache.
  5. Delivers the list to the Enrich stage.
Security Collect only accesses public http/https URLs. Internal, loopback, or private-network addresses are blocked before the request leaves the server.

Create and configure a vendor

Go to Settings → Vendors and click + Add vendor. Fill in:

FieldWhat it's for
NameIdentifies the vendor in the dashboard and in the job selector (e.g. John's Store).
Brand nameBrand applied to every product collected by this vendor. Also used by the discontinued-product check and by enrichment.
Product listing URLPaste the URL of the category/collection you want to bring in — the same one that appears in your browser when you open the product listing on the vendor's site (e.g. vendor.com/collections/dresses). This is what the platform is detected from.
Price multiplierFactor applied to the vendor's cost price to arrive at the selling price (e.g. 2.0 = double).
Sync schedule (cron)Optional. Cron expression to trigger collection automatically (e.g. 0 2 * * *).
Check for discontinued productsOptional. Daily check that marks products that disappeared from the vendor as draft in the store.

Use Test scrape to validate, Save vendor to save it. A paused vendor (is_active off) is not collected by manual jobs or by schedule.

Which URL to paste Open the vendor's category/collection you want to bring in your browser and copy the URL from the address bar. To bring in the whole store, use the broadest category URL (or /collections/all on Shopify stores). For your first collection, set a low Custom limit on the job, check the result, and only then run it in full.

Automatic platform detection

Trinis identifies the vendor's platform from the listing URL and automatically picks the right collector — you don't configure this. Platform collectors use the store's own API/JSON: they're fast and reliable. The generic collector reads the page's HTML with AI (GPT-4o) and is the fallback for non-standard sites — slower and with AI cost already at the collection stage.

Detected platformHow it collects
vtexVTEX catalog API.
shopifyPublic products endpoint of the Shopify store.
nuvemshopNuvemshop / Tiendanube API.
woocommerceWooCommerce REST API when available; otherwise, reads the store's HTML.
auto (generic)No platform recognized — AI reads the listing page and the detail pages to extract the products.

Test before saving

In the vendor form, the Test scrape button collects up to 3 products and shows title, price, EAN, image, and the detected adapter — without saving anything. Use it to confirm the URL is correct and that the vendor exposes the EAN.

Parallel collection and performance

Listing pages are fetched one at a time (pagination has to be sequential). Detail pages are fetched in parallel, up to 8 at once — which gives roughly a 5x speed gain. These values are fixed and not configurable.

Total time depends on how many products you requested and how fast the vendor's site responds. To limit the volume of a collection, use Products to sync → Custom limit when creating the job.

The cache — what gets reprocessed

Trinis keeps a cache (in Redis) per account + store + EAN, valid for 30 days. It stores each product's already-enriched description and the hash of its source image.

In the New sync form, the "Skip existing products" option is on by default. With it enabled:

Important The collection step always runs — the cache saves on AI calls (the real cost), not on downloading the vendor's pages. If the vendor is down, the collection fails even if everything were cached.

The cache expires after 30 days, aligned with the monthly usage reset, so products are naturally reprocessed every cycle. To force reprocessing before that — for example, after changing your brand's tone of voice in Settings → Enrichment — turn off "Skip existing products" when creating the job.

Scheduled collection

Fill in Sync schedule with a cron expression and Trinis will trigger that vendor's collection on its own. Examples:

If a scheduled job hits the plan's product limit, that vendor's sync is paused automatically and resumes once you have quota again (through a cycle reset, credit purchase, or upgrade). You can track each vendor's status on the Vendors tab.

Discontinued products

With "Check for discontinued products" enabled, once a day Trinis lists the EANs still sold by the vendor (without downloading detail pages) and, for products of the configured brand that disappeared, changes their status in the store to draft.

Limits and best practices

Next steps

Need help?

Couldn't collect from a vendor? Send the URL to our team.

Talk to support