Theme
Join the Waitlist

Ignyte AI · Decentralized Data Protocol

Raw data in. Proof out.

Ignyte is the machinery under a data marketplace: ingestion, PII removal before storage, quality scoring, onchain provenance — and a payment loop that pays every contributor whose row survives into a dataset.

01 — What Ignyte Is

A pipeline,
not a promise.

Most data marketplaces are a listings page with a payment button. The hard part was never the listing — it's everything that has to be true before a row is safe to sell and fair to pay for.

Ignyte does that part. Every dataset on the marketplace has been through the same five stages, and every stage leaves a verifiable trace.

Ingest

Any shape

CSV, JSONL, Parquet, archives. Schema is inferred from the file, not demanded from you before you can start.

Cleanse

PII before storage

Personal identifiers are detected and removed at the edge of the pipeline — before anything is written to a durable store.

Mint

Provenance onchain

Every surviving row carries an attributable origin. Datasets are versioned, hashed, and traceable back to the contributors who produced them.

02 — Upload

Drop
the raw.

No schema definition. No preprocessing script. Point Ignyte at a file and the pipeline takes it from there — inferring structure, flagging identifiers, and scoring quality before anything reaches the marketplace.

Drop a file or browse — CSV, JSONL, Parquet, ZIP · up to 40 GB
Queued clinical_notes_2024.jsonl 1.1 GB
Cleansing support_transcripts.csv 612 MB
Cleansed sensor_telemetry.parquet 4.8 GB

03 — Cleanse

Every row
accounted for.

customer_records_q3.csv 2.4 MB · 18,402 rows · 14 PII classes
  1. 0001 [email protected]Redacted Email
  2. 0002 +1 (415) 555-0148Redacted Phone
  3. 0003 Daniel OkaforRedacted Name
  4. 0004 4471 2299 8830 1122Redacted Card
  5. 0005 12 Rue de la Paix, 75002Redacted Address
  6. 0006 88-20-4471Redacted Gov ID
  7. 0007 [email protected]Redacted Email
  8. 0008 192.168.44.17Redacted IP
PII removed 1,204 Rows kept 17,198 Quality score 94 Provenance Minted

04 — Processing

Five stages.
No black box.

Every stage is observable and every transformation is logged against the row it touched. If a value changed, you can find out why, when, and under which rule.

01

Ingest

File is read, encoding normalised, and schema inferred from the data itself. Malformed rows are quarantined, not silently dropped.

02

Detect

Fourteen classes of personal identifier are located across every column — including free-text fields where the shape is irregular.

03

Remove

Identifiers are stripped before the row reaches a durable store. The redaction is recorded; the original value is not retained.

04

Score

Quality is measured on completeness, uniqueness, label consistency, and distribution drift against the dataset's own declared schema.

05

Mint

The cleansed dataset is versioned, hashed, and given an onchain provenance record linking it back to its contributing sources.

05 — Marketplace

Browse
the corpus.

Every listing carries its quality score, its contributor count, and its full provenance chain. Prices are set by the contributor, not by the platform. Below: example listings.

Medical Imaging 96

Anonymized chest radiograph set

Rows 1,240,000 Contributors 12 Format DICOM → Parquet
420IGN
Conversational 91

Customer support transcripts

Rows 840,000 Contributors 34 Format JSONL
285IGN
Financial 94

Redacted ledger event stream

Rows 3,400,000 Contributors 8 Format Parquet
610IGN
Geospatial 88

Road surface imagery, 14 regions

Rows 620,000 Contributors 21 Format COCO → Parquet
340IGN
Linguistic 92

Code-switched speech corpus

Rows 410,000 Contributors 57 Format WebAudio → Parquet
195IGN
Industrial 90

Multi-site sensor telemetry

Rows 7,800,000 Contributors 15 Format Parquet
880IGN

06 — Token & Payment Loop

Paid on use.
Not on upload.

A contributor is paid every time their rows are consumed — not once, at the moment of upload. Revenue from every acquisition is split across the provenance graph, weighted by how much of the delivered dataset each source accounts for.

  1. Upload Contributor submits raw data.
  2. Cleanse PII removed, quality scored.
  3. Mint Provenance recorded onchain.
  4. Consume Buyer acquires the dataset.
  5. Split Payment distributed per row.
Revenue split → back to source

07 — Contributor Earnings

Your rows.
Your revenue.

Earnings accrue per row consumed, not per upload. A dataset that keeps getting acquired keeps paying the people who made it — for as long as it stays in use.

0.4%
Platform fee per acquisition
Per-row
Payout granularity
Onchain
Settlement, publicly verifiable
Medical Imaging 92
Conversational 74
Industrial 61
Geospatial 47
Linguistic 33

08 — Trust

Verify,
don't trust.

The pipeline's guarantees are structural, not contractual. You don't have to believe the platform — you can check the chain, check the code, and check that the raw values are gone.

01

PII never lands

Identifiers are removed at the edge of the pipeline, before the row is written to a durable store. There is no "raw" tier holding the original values.

02

Provenance is public

Every dataset carries a hash and an onchain record of its contributing sources. Anyone can independently reconstruct where a row came from.

03

Payouts are auditable

Revenue splits are computed from the provenance graph and settled onchain. Contributors can verify their own share without asking the platform for it.

09 — Waitlist

Bring the
raw data.

We're onboarding contributors and data buyers in cohorts. Leave an address and we'll let you know when your cohort opens.