Ignyte AI · Decentralized Data Protocol
Raw data in. Proof out.
01 — What Ignyte Is
A pipeline,
not a promise.
Most data marketplaces are a listings page with a payment button. The hard part was never the listing — it's everything that has to be true before a row is safe to sell and fair to pay for.
Ignyte does that part. Every dataset on the marketplace has been through the same five stages, and every stage leaves a verifiable trace.
Any shape
CSV, JSONL, Parquet, archives. Schema is inferred from the file, not demanded from you before you can start.
PII before storage
Personal identifiers are detected and removed at the edge of the pipeline — before anything is written to a durable store.
Provenance onchain
Every surviving row carries an attributable origin. Datasets are versioned, hashed, and traceable back to the contributors who produced them.
02 — Upload
Drop
the raw.
No schema definition. No preprocessing script. Point Ignyte at a file and the pipeline takes it from there — inferring structure, flagging identifiers, and scoring quality before anything reaches the marketplace.
03 — Cleanse
Every row
accounted for.
- 0001 [email protected]Redacted Email
- 0002 +1 (415) 555-0148Redacted Phone
- 0003 Daniel OkaforRedacted Name
- 0004 4471 2299 8830 1122Redacted Card
- 0005 12 Rue de la Paix, 75002Redacted Address
- 0006 88-20-4471Redacted Gov ID
- 0007 [email protected]Redacted Email
- 0008 192.168.44.17Redacted IP
04 — Processing
Five stages.
No black box.
Every stage is observable and every transformation is logged against the row it touched. If a value changed, you can find out why, when, and under which rule.
Ingest
File is read, encoding normalised, and schema inferred from the data itself. Malformed rows are quarantined, not silently dropped.
Detect
Fourteen classes of personal identifier are located across every column — including free-text fields where the shape is irregular.
Remove
Identifiers are stripped before the row reaches a durable store. The redaction is recorded; the original value is not retained.
Score
Quality is measured on completeness, uniqueness, label consistency, and distribution drift against the dataset's own declared schema.
Mint
The cleansed dataset is versioned, hashed, and given an onchain provenance record linking it back to its contributing sources.
05 — Marketplace
Browse
the corpus.
Every listing carries its quality score, its contributor count, and its full provenance chain. Prices are set by the contributor, not by the platform. Below: example listings.
Customer support transcripts
Redacted ledger event stream
Road surface imagery, 14 regions
Code-switched speech corpus
Multi-site sensor telemetry
06 — Token & Payment Loop
Paid on use.
Not on upload.
A contributor is paid every time their rows are consumed — not once, at the moment of upload. Revenue from every acquisition is split across the provenance graph, weighted by how much of the delivered dataset each source accounts for.
- Upload Contributor submits raw data.
- Cleanse PII removed, quality scored.
- Mint Provenance recorded onchain.
- Consume Buyer acquires the dataset.
- Split Payment distributed per row.
07 — Contributor Earnings
Your rows.
Your revenue.
Earnings accrue per row consumed, not per upload. A dataset that keeps getting acquired keeps paying the people who made it — for as long as it stays in use.
08 — Trust
Verify,
don't trust.
The pipeline's guarantees are structural, not contractual. You don't have to believe the platform — you can check the chain, check the code, and check that the raw values are gone.
PII never lands
Identifiers are removed at the edge of the pipeline, before the row is written to a durable store. There is no "raw" tier holding the original values.
Provenance is public
Every dataset carries a hash and an onchain record of its contributing sources. Anyone can independently reconstruct where a row came from.
Payouts are auditable
Revenue splits are computed from the provenance graph and settled onchain. Contributors can verify their own share without asking the platform for it.
09 — Waitlist
Bring the
raw data.
We're onboarding contributors and data buyers in cohorts. Leave an address and we'll let you know when your cohort opens.