Gapfill — building-product attributes the supplier never sent
An enrichment layer for building-product catalogues: pull supplier data from websites, inventory systems and PDFs, then fill the fields the supplier never sent. Five research passes refuted the original moat — vision-enrichment vendors DO serve building materials, including an Australian one — but confirmed a narrow documented gap. Full measurements now live in four linked research annexes. The binding constraint is no longer capability, it is whether anyone pays.
researchingecommerce / product-data-automationfree idea
Supplier product data arrives as a mess and has to leave as a catalogue. This is the research record for an idea about closing that gap for building products — tiles and lighting first — and it is mostly a record of the idea getting harder, not easier. Three proposed moats were researched and all three collapsed. The claim the page leaned on hardest was refuted by its own evidence. What survives is narrower, better evidenced, and still gated on one question nobody has answered.
Research & evidence
Showing 20 of 42 entries behind this idea, oldest first. Open one to read it in full.
Evidence 15
Sources, findings, and competitor scans.
evidenceCOMPETITOR SCAN — the category is occupied at all three price tiers. This idea is not novel.
COMPETITOR SCAN — the category is occupied at all three price tiers. This idea is not novel.
TIER 1, ENTERPRISE PIM WITH A BUILDING-MATERIALS VERTICAL ALREADY BUILT. Pimberly runs dedicated "Building Materials Manufacturers" and "PIM for Construction" pages. Sales Layer runs "PIM for the Building Materials industry" and explicitly markets handling of building-materials attributes — dimensions, density, fire/water resistance, SKUs — plus DAM for technical layouts, 3D models and certifications. Bluestone PIM and Apimio publish construction/building-materials content too. Pricing anchor: Pimberly starts around USD 30,000/year (Capterra). So "building materials" is not an unserved vertical; it is a named, marketed segment.
TIER 2, AI EXTRACTION + SUPPLIER ONBOARDING. SKU Launch is close to the idea as written: pulls structured values from PDF spec sheets including tables, bullet lists and multi-column layouts mapped to your attribute schema; reads specs off product images; runs web-search agents to find authoritative product data; and replaces spreadsheet collection with a portal where suppliers confirm pre-filled data (claims completion rising from ~40% to 80%+). Also in this tier: dataX.ai, Klizer, Hypotenuse AI, getcatalog.ai.
TIER 3, SMB SHOPIFY APPS DOING THE EXACT DESCRIBED WORKFLOW. Shopif-AI takes a supplier PDF catalogue up to 300 pages and extracts names, prices, variants, images and specs, then publishes to Shopify. Importier imports CSV/Excel/PDF, auto-detects headers and auto-maps 40+ Shopify fields. Apport does AI bulk PDF/CSV import. Automated Commerce (AI PIM) imports from supplier CSVs, PDFs and ERPs and generates metafield values in bulk. Matrixify documents a Claude-via-MCP workflow that maps any supplier file into a Shopify import.
CONCLUSION: the "200 products by hand into Shopify" pain is real enough that a competitive market already serves it end to end, from free-ish apps to 30k/year platforms.
evidenceTHE CANONICAL SCHEMA ALREADY EXISTS AS AN OPEN STANDARD — this kills the primary proposed moat.
THE CANONICAL SCHEMA ALREADY EXISTS AS AN OPEN STANDARD — this kills the primary proposed moat.
The page claimed the defensible asset was "a maintained, opinionated taxonomy of tile/lighting spec fields with the synonym maps that make it work". That taxonomy exists, is international, and is free.
ETIM (Electrotechnical Information Model) is the open international classification standard for technical products. It defines classes, codes and standardised technical attributes precisely so manufacturers, distributors and retailers can exchange product data unambiguously. Its stated scope is electrical, HVAC, plumbing AND building materials — the exact territory of this idea. Lighting sits at its historical core, since it began as an electrotechnical standard.
BMEcat is the paired XML exchange format: manufacturers structure catalogues as BMEcat with ETIM classification embedded, giving a standard machine-readable route from supplier to distributor. A published ETIM BMEcat guideline exists. Sources also note ETIM is increasingly the feature-data backbone for BIM objects, linking specify-procure-install.
For ceramic tiles, ISO 13006:2018 already defines classification and characteristics — water absorption groups, extruded vs dry-pressed, dimensional tolerances, surface quality — plus canonical terminology for glaze, engobed and rectified. That is a substantial share of the "14-20 fields" already standardised.
CAVEAT AND THE ONE REMAINING OPENING: ETIM coverage is strongest in electrical/HVAC/plumbing. Searches did not confirm class depth for ceramic tiles and flooring, and sources explicitly noted flooring coverage could not be established. Tiles may be genuinely thin even though lighting is well covered. Worth checking directly against the ETIM class browser — but a gap in an open standard is a standards-body problem, not obviously a startup.
NET: do not build a proprietary taxonomy. Adopt ETIM plus ISO 13006 and compete on mapping INTO them, not on owning them.
evidenceGAP-FILLING RESEARCH — inference of missing attributes is shipped, but not in this vertical.
GAP-FILLING RESEARCH — inference of missing attributes is shipped, but not in this vertical.
Vision-based attribute enrichment is real and in market: Stylitics, Velou, Constructor, Feedonomics, Hypotenuse and Envive combine computer vision with NLP to tag visual attributes and fill incomplete vendor data. Documented techniques include RAG that infers likely values from similar products and category norms, and multimodal passes that separate explicit attributes (stated) from implicit ones (inferred from context and visual cues).
CRUCIAL NUANCE: the cited application industries are fashion, electronics and automotive, and the canonical example attributes are neckline, silhouette and print. Building materials are not named. So the CAPABILITY is commodity, but its APPLICATION to tile and lighting spec vocabulary is not evidently occupied — unlike extraction, which is occupied all the way down to cheap Shopify apps.
DETERMINISTIC DERIVATION IS ALSO REAL, and the standard supplies the rules. ISO 13006 ties water absorption to group: <0.5% = BIa, 0.5-3% = BIb, 3-6% = BIIa, 6-10% = BIIb, >10% = BIII. Porcelain is definitionally water absorption below 0.5%, i.e. "impervious". So porcelain-versus-ceramic is DERIVED, never asked for.
MISSING FIELDS CARRY SIGNAL. PEI grades the wear resistance of a GLAZED surface; unglazed tiles, including many through-body porcelains, do not receive a PEI rating at all. An absent PEI is therefore evidence of an unglazed or through-body product, not merely a hole in the data.
DO NOT OVER-DERIVE. Sources treat water absorption, PEI and COF/slip resistance as distinct complementary attributes. Slip rating is a measured test result and is NOT reliably derivable from finish or appearance — which is exactly why the liability risk on this page stays live.
evidenceCORRECTION TO THE COMPETITOR SCAN — Matrixify was overstated as a competitor. From direct user knowledge: Matrixify provides NO user interface into t…
CORRECTION TO THE COMPETITOR SCAN — Matrixify was overstated as a competitor. From direct user knowledge: Matrixify provides NO user interface into the data. Its configuration is over CSV/Excel only. It is a bulk import/export engine, not a place where product data lives and can be inspected or corrected.
WHY THIS IS STRUCTURAL, NOT A MISSING FEATURE. The surviving thesis on this page is inference: filling fields the supplier never sent, some derived by rule and some inferred from imagery. An inferred value is epistemically different from an extracted one — it needs a confidence score, a source pointer, and a human confirm/reject before it is safe to publish. A CSV round-trip cannot express "this colour family was inferred from the product image at moderate confidence, click to confirm". The moment the artefact is a spreadsheet, provenance and confidence collapse into flat cell values and the review step becomes unauditable.
So file-based tools are not merely behind on UI — they are ARCHITECTURALLY EXCLUDED from the inference product. This meaningfully narrows the competitor set previously recorded.
DO NOT OVERCLAIM THOUGH. Some app-tier competitors do have review surfaces: Importier is documented as generating descriptions and then letting you review everything in a bulk editor. The honest distinction is not "UI vs no UI" but PERSISTENT CANONICAL RECORD vs IMPORT BATCH. A bulk editor reviews one import into one platform and is then discarded. What the inference thesis requires is a durable record where every field carries extracted/derived/inferred/human-asserted status, corrections persist, and those corrections accumulate into the labelled corpus identified as the only defensible asset.
CONSEQUENCE FOR POSITIONING: the review interface is not a nice-to-have layered on top of the pipeline. It IS the product surface, because it is simultaneously the safety mechanism for the never-infer compliance rule and the collection mechanism for the corpus.
evidenceDEEP MARKET SCAN — two findings that were missed by the first pass.
DEEP MARKET SCAN — two findings that were missed by the first pass.
1. A DIRECT VERTICAL COMPETITOR EXISTS. MindHarbor serves "Tile, Flooring & Building Materials Distribution" specifically and states it normalises vendor price lists and product catalogues, plus supplier price and product list management, multi-UOM handling and variant management, delivered as ERP integration and "Digital Intelligence". That is this idea's core value proposition, already named, already sold into this exact vertical. It appears positioned as an ERP integration services firm rather than a self-serve product, which is a meaningful difference in shape but not in claim.
The surrounding tier is also crowded with vertical ERP: Epicor runs a dedicated Tile Distribution solution; Klipboard ERP covers flooring and surfaces including supplier catalogue integration and multiple price lists; Comp-U-Floor automatically exchanges price catalogues, POs and vendor invoices between manufacturers, distributors and flooring retailers; Acctivate does tile distribution on QuickBooks. So catalogue and price-list ingestion is table stakes in flooring ERP.
2. THE SPEC-LIBRARY CHANNEL IS CONSOLIDATING, WHICH DAMAGES PIVOT A. Anguleris acquired Concora Spec in February 2025, adding it to a portfolio that already held BIMsmith (product research and BIM content), Swatchbox (physical sample fulfilment, integrated with BIMsmith so manufacturers can offer samples inside the BIM design workflow) and Modlar. Pivot A assumed a fragmented set of independent spec libraries to syndicate into. Instead one owner now holds research, content, sampling and specification. That means fewer integration targets, a stronger counterparty, and an incumbent with obvious motive to build ingestion itself.
NET: the vertical is not empty, and the escape hatch is narrower than the previous entry claimed. Both need weighing before any build decision.
evidenceOPEN-SOURCE STACK — nearly every component exists, which cuts both ways.
OPEN-SOURCE STACK — nearly every component exists, which cuts both ways.
TAXONOMY, THE BIG ONE. The ETIM Classification Model is published under the Open Data Commons Attribution Licence, free for everyone. Downloadable full-model files plus a REST API (etimapi.etim-international.com), Swagger-documented, OAuth2 with client_id/secret, offering Release and Details services with filtering and paging. The canonical schema is not merely public — it is programmatically consumable today at zero licence cost.
PIM BACKBONE. Pimcore is open source and self-hostable, spanning PIM, DAM and MDM, scoring higher than Akeneo on data modelling and governance. CRITICAL: Akeneo Community Edition has had no new features since 2023, last major CE release v7 March 2023, support only to September 2026 — development moved to paid editions. A live gap. Alternatives: AtroPIM, UnoPIM, PixeePIM. Cost check: running Akeneo CE cited at EUR 1,500-3,000/month all-in; Pimcore production needs a PHP/Symfony developer.
DOCUMENT EXTRACTION. Docling hit 97.9% on complex tables but is slow. Marker uses a 5-stage pipeline of small specialised models — best local option, faster than Docling. LlamaParse is fastest and cleanest but cloud-hosted, so PDFs leave your infrastructure. Unstructured has the strongest OCR.
ATTRIBUTE EXTRACTION, DIRECTLY ON-POINT. wbsg-uni-mannheim/ExtractGPT covers LLM attribute value extraction; wbsg-uni-mannheim/wdc-pave covers extraction AND NORMALIZATION — the exact two-step needed. google/langextract does structured extraction with PRECISE SOURCE GROUNDING, this page's provenance requirement. Structured output: Instructor, Outlines, PydanticAI, Guidance.
DEDUPLICATION. Splink does probabilistic record linkage — a million records on a laptop in about a minute, 100M+ on Spark. Also Zingg and dedupe.
IMPLICATION: build cost is assembly, not invention — cheap for us, equally cheap for anyone else. Reinforces that the corpus, not the pipeline, is the only asset.
evidenceRESEARCH LOG 3/9 — SOURCE INDEX A: PIM AND VERTICAL SOFTWARE.
RESEARCH LOG 8/10 — RESEARCHED BUT NOT PREVIOUSLY RECORDED, PART 1. (Log extended to ten parts.)
PIM ECONOMICS, FULLER PICTURE. Akeneo Community Edition is free to licence, but budget EUR 15,000-40,000 for initial setup via a partner, and an operational instance runs EUR 1,500-3,000 per month all-in (server, maintenance, plugins), excluding in-house developer time. Pimcore commercial deployments start at Professional, USD 9,900/yr. Pimcore Community has no cloud-hosted option — self-host, and you need a PHP/Symfony developer to configure the server and build connector logic per channel.
CAPABILITY SCORING. One comparison scores Pimcore 35/40 across eight PIM capabilities against Akeneo 31/40, Pimcore leading on data modelling and governance, digital asset linkage, localisation and interoperability.
MATERIAL OMISSION NOW CORRECTED: AKENEO LEADS ON SUPPLIER ONBOARDING, via its Supplier Data Manager. That is directly adjacent to this idea's ingest half and was missed in the earlier competitor scan. Akeneo competes on supplier onboarding specifically, not merely as a generic PIM.
THE CONNECTOR / iPaaS LAYER WAS ALSO MISSED. Stacksync markets a Shopify-and-Cin7 integration, stating both standard and custom objects can be synchronised. Make.com offers Cin7-to-WooCommerce automation. Beyond Cin7's native connectors there is a generic iPaaS tier able to move custom fields between systems with no purpose-built product. This weakens the multi-platform-publishing claim further than previously recorded.
CIN7 FIELD DEPTH, PRECISE. Cin7 Core syncs products to and from Magento and Shopify. Cin7 Omni exposes an Update Recently Modified Products Field Mapping page mapping Omni fields to Shopify fields, Omni data replacing Shopify data on upload. But for WooCommerce only stock availability syncs automatically — descriptions and other details need manual sync. Field depth is uneven per platform.
evidenceRESEARCH LOG 9/10 — RESEARCHED BUT NOT PREVIOUSLY RECORDED, PART 2: TECHNICAL.
RESEARCH LOG 9/10 — RESEARCHED BUT NOT PREVIOUSLY RECORDED, PART 2: TECHNICAL.
PDF PARSER BENCHMARK, PRECISE. Same test: LlamaParse 8.54 seconds, Marker 47 seconds, Docling 1 minute 50 seconds. The architectural reason matters — Docling runs one large vision-language model reading each page as an image and generating table structure token by token, inherently slow; Marker runs a 5-stage pipeline of smaller specialised models doing mostly classification and detection, avoiding token-by-token generation. Docling was separately reported at 97.9% accuracy on complex table extraction in sustainability reports. Unstructured is noted for strongest OCR.
CRITICAL CONSTRAINT NOT PREVIOUSLY STATED: LlamaParse is cloud-hosted on LlamaCloud, so documents leave your infrastructure. Supplier price lists routinely carry confidential trade pricing, which disqualifies the fastest option and pushes the build toward Marker or Docling running locally. This is an architectural fork, not a preference.
OFFICE.JS CONSTRAINTS, PRECISE. Add-ins are web apps in a sandboxed iframe running in Office on the web, Windows, Mac and iPad, deployable centrally by admins. They read and write worksheets, ranges, tables and charts, extend the ribbon and context menu, add task panes, and add custom functions callable from cells. But they are strictly sandboxed by the browser engine and CANNOT ACCESS THE LOCAL DRIVE — so bulk ingestion of local supplier files from inside an add-in is constrained. This reinforces Excel as an import/export surface rather than an ingestion path.
ETIM SCOPE EXCLUSIONS, EXPLICIT. Products outside electrical, HVAC, plumbing and building materials — consumer electronics, food, fashion, industrial machinery — have limited or no ETIM coverage. ETIM is the right standard for this vertical, and would not generalise if the idea expanded beyond building products.
evidenceRESEARCH LOG 10/10 — RESEARCHED BUT NOT PREVIOUSLY RECORDED, PART 3: COMMERCIAL.
RESEARCH LOG 10/10 — RESEARCHED BUT NOT PREVIOUSLY RECORDED, PART 3: COMMERCIAL.
SHOPIFY B2B GAP. Shopify lacks native functionality for sales reps to take orders at trade shows or in the field — a documented gap versus purpose-built B2B wholesale platforms. For a building-products supplier selling to architects and trade, the storefront does not cover how the sale actually happens. The storefront may not be the centre of this customer's world, which cuts against any Shopify-first framing.
SHOPIFY APP BILLING, COMMERCIALLY MATERIAL. A Shopify developer-community thread covers off-platform billing for connector apps where the vendor is a standalone SaaS acquiring zero merchants through Shopify, and whether an exemption applies versus split billing. If the Shopify app is only a publish channel onto a standalone product — the recommended architecture — billing needs deliberate handling and is not automatic. Unresolved; flagged for whoever builds.
SUPPLIER-CONFIRMATION PATTERN, WORTH STEALING. SKU Launch replaces spreadsheet collection with a portal where suppliers CONFIRM PRE-FILLED DATA rather than typing, claiming completion rises from about 40% to over 80%. Same confirm-not-type principle this page reached independently for the review queue — evidence the pattern works, and that a competitor already applies it supplier-side.
TCNA SCALE. The Tile Council of North America has over 200 member companies representing roughly 90% of ceramic tile production in North America, and publishes the 2026 TCNA Handbook. A body with that concentration is a plausible distribution partner or a plausible source of a competing data initiative.
BIM LIBRARY SCALE. BIMobject provides tens of thousands of product families; NBS National BIM Library offers generic and manufacturer-specific objects; ARCAT serves architects, spec writers and contractors. The market is described as highly fragmented — consistent with the Anguleris consolidation already recorded.
evidenceRESEARCH LOG 1/10 — METHOD. (Re-filed as evidence: originally posted as a comment, which meant it rendered in the comment thread instead of the resea…
RESEARCH LOG 1/10 — METHOD. (Re-filed as evidence: originally posted as a comment, which meant it rendered in the comment thread instead of the research record.)
All research on this page was done on 29 July 2026 by web search only. No primary interviews, no vendor demos, no trials, no pricing calls, no product testing. Every claim is desk research. One exception: the Matrixify correction came from the idea owner's direct product knowledge.
23 QUERIES RUN, grouped.
COMPETITORS (7): PIM for building materials and tile suppliers; AI extraction of product data from supplier PDFs into ecommerce; SKU Launch extraction pricing and competitors; Pimberly and Sales Layer building-materials PIM pricing; Shopify apps for bulk supplier-catalogue import; tile and flooring distributor catalogue management software; MindHarbor normalising vendor price lists.
STANDARDS (4): ETIM classification and BMEcat; ETIM class coverage for tiles and flooring plus scope limits; ETIM API licence and developer access; TCNA and ISO 13006 ceramic tile attributes.
INFERENCE (2): AI inference of missing product attributes from imagery; ISO 13006 water absorption in relation to porcelain, rectified, PEI and slip.
CHANNELS (2): architect specification platforms BIMobject, ARCAT, NBS Source, Concora; building-products data syndication including Concora and Swatchbox.
INTEGRATION (1): Cin7 product field sync to Shopify, WooCommerce and Magento.
OPEN SOURCE (4): open-source PIM Pimcore versus Akeneo; PDF table extraction Docling, Marker, LlamaParse; entity resolution Splink and dedupe; product attribute extraction with LLM structured output.
FORM FACTOR (3): Shopify app store versus standalone SaaS; Excel add-in Office.js versus web app; Shopify metafields and metaobjects bulk operation limits.
Risks 1
Reasons this could fail.
riskVERDICT AFTER RESEARCH: the pain is real and validated, but all three proposed moats are occupied.
VERDICT AFTER RESEARCH: the pain is real and validated, but all three proposed moats are occupied.
Moat 1, the building-materials taxonomy — GONE. ETIM is an open international standard covering building materials; ISO 13006 already standardises ceramic tile classification. Owning a taxonomy is not available; it is a free public good.
Moat 2, AI extraction from messy supplier formats — GONE. Commodity capability shipping in cheap Shopify apps (Shopif-AI, Importier, Apport) and mid-market platforms (SKU Launch). LLM improvement makes it cheaper every quarter, compressing margin rather than building advantage.
Moat 3, multi-platform publishing — ALSO GONE. Cin7 Core documents Shopify, Magento and WooCommerce integrations that sync products; Cin7 Omni exposes editable Cin7-to-Shopify field mappings. Generic PIMs treat channel syndication as table stakes.
THE SQUEEZE: this sits in the valley between a cheap single-platform Shopify importer and a roughly USD 30k/year enterprise PIM. The buyer there — a showroom with a few thousand SKUs across twenty suppliers and two storefronts — is real but small, hard to reach, and low willingness to pay. Both sides are expanding into it: Shopify apps adding channels and PIM features, PIM vendors adding self-serve tiers.
HONEST READ: as venture-scale software, weak. The differentiated work remaining is integration and mapping labour — a services shape, not a licence shape. Test the service framing first: cheaper, monetises immediately, and produces the supplier-format corpus that would be the only compounding asset if a product version ever justified itself.
CAVEAT BEFORE WRITING IT OFF: a Cin7 source notes WooCommerce syncs stock automatically but descriptions and other product details need manual sync. If shallow field coverage is typical across these integrations, the enrichment gap is wider than this verdict assumes. Verify before killing.
Pivots 1
Alternative shapes worth testing.
pivotTWO PIVOTS THAT SURVIVE THE COMPETITOR SCAN. Both start from one observation: the ecommerce channel is saturated, the ARCHITECT channel is not — and…
TWO PIVOTS THAT SURVIVE THE COMPETITOR SCAN. Both start from one observation: the ecommerce channel is saturated, the ARCHITECT channel is not — and the original context said these suppliers sell to architects.
PIVOT A — PUBLISH TO THE SPECIFICATION CHANNEL, NOT JUST THE STORE. Architects do not shop a Shopify storefront; they specify from BIM and specification libraries. Manufacturers submit products to BIMobject, ARCAT, NBS Source and Concora, and the BIM-objects market is documented as highly fragmented. ETIM is separately described as becoming the standardised feature data behind BIM product objects, linking specify-procure-install. So the differentiated pipeline is: normalise once against ETIM/ISO 13006, then publish to storefronts AND to the specification libraries where the actual buyers are. No Shopify importer does this — they are ecommerce-only by construction. Generic PIMs syndicate to retail channels, not spec libraries. This is the one place the scan found genuine white space. Unverified and load-bearing: whether suppliers feel enough pain getting onto those libraries to pay, and whether the libraries permit programmatic submission.
PIVOT B — SELL THE SERVICE, NOT THE SOFTWARE. Given the origin was a digital marketing strategy conversation, the fastest honest monetisation is a productised data-onboarding service for building-product suppliers, delivered with off-the-shelf tools rather than a new platform: take the supplier's website, Cin7 export and PDFs, produce a clean normalised catalogue, publish to their storefronts, charge per catalogue or per SKU. Zero build risk, immediate revenue, and it directly answers the questions the page cannot — real cadence, real willingness to pay, how many storefronts, how clean the SKU join is.
SEQUENCING: run B to fund and inform A. B is a business this week; A is the only version with a defensible product story, and B is the cheapest way to earn the right to build it.
Proposed refinements 3
Proposals for the document above. Open ones are awaiting merge.
Section: design
Proposal:
Ingest is three source classes carrying different halves of the record.
1. SUPPLIER WEBSITE (crawl). Structured HTML spec tables plus images. Cheapest, needs no supplier cooperation, and uniquely enables continuous monitoring — re-crawl to catch new products, discontinued lines, spec revisions. Costs: ToS, brittleness, politeness limits.
2. INVENTORY / ERP (Cin7 Core or Omni, Unleashed, DEAR, NetSuite). Clean and API-accessible, but carries COMMERCIAL fields: SKU, barcode, cost, price, stock, supplier. Not the 14-20 architectural spec fields — nobody types slip rating or PEI class into an inventory system.
3. PDF / EXCEL (price lists, spec sheets, brochures). Messiest, highest extraction cost, but where spec data actually lives: thickness, finish, slip rating, frost resistance, rectified edge, application.
CONSEQUENCE: the canonical record is a MERGE, not an import. Commercial truth from inventory, spec truth from documents and website, joined on SKU or supplier code — which will not always match cleanly. Merge policy becomes first-class: per-field source-of-truth rules, conflict resolution when Cin7 and the PDF disagree, provenance on every field.
COMPETITIVE CONSEQUENCE: Cin7 reportedly already publishes to Shopify, WooCommerce and Magento (verify). For a Cin7 customer, publish-many ALREADY EXISTS and is paid for — removing publishing as a reason to buy. What Cin7 cannot do is carry normalised spec metadata. Honest positioning may be a SPEC LAYER ENRICHING an existing sync, not replacing it.
Rationale:
The owner specified three ingest sources: supplier websites, inventory databases such as Cin7, and PDF/Excel files. These are not variations of one pipeline — inventory systems hold commercial fields while spec data lives in documents, making the canonical record a merge with per-field source-of-truth rules and provenance. It also surfaces a competitive fact the page missed: Cin7 already syncs to all three target platforms, so publis
Section: design
Proposal:
Suppliers omit fields, and missing values get filled by deduction or by an experienced person looking at the product. Three mechanisms, different trust levels, never to be blurred.
1. DERIVED — rule from standard. ISO 13006 maps water absorption to group to porcelain/ceramic; absent PEI implies unglazed or through-body. Deterministic and auditable.
2. INFERRED — vision or category norms. Colour family, look and style (wood-look, concrete-look, terrazzo), texture, pattern, room suitability. Probabilistic. The tacit industry judgment.
3. ASKED — chase the supplier. The only route for anything unsafe to guess.
HARD RULE, SPLIT THE FIELDS IN TWO:
COMPLIANCE/SAFETY (slip COF or R-value, fire rating, frost resistance, load ratings, PEI) are MEASURED TEST RESULTS. Never infer, never publish a guess — leave blank and chase. A wrong value here is the liability event already on the risk register.
MERCHANDISING/DISCOVERY (colour family, look, style, texture, room, finish family, application tags) can be inferred freely. Wrong costs a slightly worse filter.
THE REFRAME: merchandising fields are exactly what suppliers never provide and stores most need for filtering, SEO and comparison. Inference targets precisely the gap, landing this back in DIGITAL MARKETING value rather than data entry.
THE MOAT: every human correction in the review queue is a labelled example linking image plus partial fields to a correct value, in a vertical the enrichment vendors serve fashion instead of. The taxonomy is public and extraction is commodity; this corpus is neither.
Rationale:
The owner noted that missing supplier data is deduced from other fields or from a human looking at the product with industry experience. This is a materially different capability from extraction — extraction is commodity and occupied, whereas inference over building-materials vocabulary is not evidently served, since the vision-enrichment vendors target fashion, electronics and automotiv
Section: prototype
Proposal:
FORM FACTOR: standalone web app at the core, thin connectors at the edges. Not a Shopify app, not an Excel plug-in.
Shopify's own guidance names four cases where standalone beats embedded, and ALL FOUR apply: the UI needs more real estate than admin-chrome width (review queue = image, source snippet, confidence, canonical value, side by side); it is used by non-merchant roles (suppliers confirming data, agency staff, architects); it spans multiple storefronts; and it is a full tool where Shopify is one integration among many.
Accept the cost knowingly: installed-to-first-use runs 70-85% embedded versus 45-60% standalone. So a Shopify app should still exist — as a DISTRIBUTION and publish channel, a thin client onto the standalone record, not the product.
EXCEL IS THE WRONG HOME, RIGHT DOORWAY. Office.js add-ins are web apps in a sandboxed iframe across web/Windows/Mac/iPad, so it is viable technically — but a spreadsheet cannot carry per-field confidence, provenance or extracted/derived/inferred status, cannot support multi-user review, and does not persist corrections. Same structural disqualification already established for Matrixify. Excel and Sheets belong as IMPORT/EXPORT surfaces only.
PUBLISH PATH: Shopify Bulk Operations API is right — async GraphQL returning JSONL, tens of thousands of records without hitting standard rate limits, though one bulk operation per type runs at a time per shop and mutations can partially fail, so retry logic is ours. Watch 128 metaobject definitions per app, 30 fields per definition.
Rationale:
The owner asked what form the solution should take. The answer follows from the page's own conclusion that the review surface IS the product: only a persistent multi-user record with per-field provenance can enforce the never-infer compliance rule and accumulate the correction corpus. Shopify's published embedded-versus-standalone criteria all point the same way, and Office.js carries the same structural
Loading comments...