Data
For data partners
Upload your data, publish it on the marketplace, and earn 70% of what other organisations pay when their searches return it. Each month TRIVDA issues a statement from the ledger, and you invoice TRIVDA for it.
1. Upload
Upload in the dashboard under Add data, or with the API:
curl -sS "$TRIVDA_URL/api/v1/datasets/upload" \
-H "Authorization: Bearer $TRIVDA_API_KEY" \
-F "file=@handbook.pdf" \
-F "dataset_id=labour-law-handbook" \
-F "data_owner=Example Förlag AB" \
-F "license=CC-BY-4.0" \
-F "granted_uses=search,rag" \
-F "jurisdiction=SE"
| Field | Meaning |
|---|---|
file required | One file of up to 20 MB: PDF, Office documents, spreadsheets, HTML, Markdown, plain text, images (OCR), audio and more. A .zip becomes one document per file inside it. |
dataset_id | Your own name for the collection to add to. The first upload under a name creates the collection, with a new id (col_…, in the response as dataset_id) and this form's licence fields as its terms; later uploads under the same name add to it. Names are your organisation's own: another organisation using the same name gets a collection of its own. Defaults to a slug of the file name. |
collection_id | Instead of dataset_id: a collection you made with POST /api/v1/collections. Its terms stay as they are. |
doc_key | The document's name inside the collection, usually its path in your folder (Avtal/2024/lease.pdf). Default: the file name. Uploading new bytes under the same doc_key replaces the previous version. |
data_owner | Who holds the rights. |
license | The SPDX id of the licence, default CC-BY-4.0, or a LicenseRef- id for your own terms. |
granted_uses | Comma-separated uses buyers may make: search, rag, and so on. Default search,rag. |
jurisdiction | Where the rights were granted. Default global. |
offered | true publishes on upload. Default false: private. |
upload_id | Your own id for the upload (letters, digits, - and _, up to 64), so you can follow it with GET /api/v1/uploads/{upload_id} and stop it at any stage with POST /api/v1/uploads/{upload_id}/cancel. |
Bytes your organisation already holds, in any collection, are not processed again: the response says duplicate: true and points at the collection that holds them.
By uploading you state that you hold the rights to license the data on these terms. Each file is parsed, cleaned, checked for personal data, split into passages and run through the quality gate; only passages that pass are indexed. The response counts indexed and quarantined passages and gives the readiness tiers.
A large file can take longer than your HTTP client waits. If the request times out, do not send the file again: ask GET /api/v1/uploads/{upload_id} until its state is no longer running. It ends as done (with the indexed and quarantined counts), failed (with an error) or cancelled; unknown means the server holds no record of it, for example an upload that finished more than an hour ago.
2. Publish
An upload is private: only your organisation can search it. Publish a dataset when you want other organisations to find it:
curl -sS -X POST "$TRIVDA_URL/api/v1/datasets/labour-law-handbook/offer" \
-H "Authorization: Bearer $TRIVDA_API_KEY"
curl -sS -X POST "$TRIVDA_URL/api/v1/datasets/labour-law-handbook/withdraw" \
-H "Authorization: Bearer $TRIVDA_API_KEY"
- Once offered, every paying organisation can search the dataset. Buyers see your licence, attribution duties and source with every passage; they never see your other data.
- Withdraw takes it off the marketplace at once. It stays indexed and searchable by you. Results that other organisations had cached from it are dropped.
- Delete (
POST /api/v1/datasets/{id}/erase?reason=owner_removed) removes its passages, the stored originals and every cached result that quotes it. You can upload the same data again later.
Collections and documents
A collection is the unit you publish, license and earn on: a folder of thousands of files is one collection with one set of terms, not thousands of datasets. Its id is also its dataset_id, so offer, withdraw, delete and earnings take it.
POST /api/v1/collectionsmakes an empty collection with its terms;GET /api/v1/collectionslists yours;PATCH /api/v1/collections/{collection_id}renames one.GET /api/v1/collections/{collection_id}/documentsandGET /api/v1/collections/{collection_id}/folderspage through what it holds;GET /api/v1/collections/{collection_id}/documents/{doc_id}shows one document's versions.DELETE /api/v1/collections/{collection_id}/documents/{doc_id}removes one document (its passages, original and cached results); the rest of the collection stays.POST /api/v1/collection-rulessends new uploads whosedoc_keymatches a pattern (Protokoll/*) to another collection.
Coming soon: a data card and listing terms per collection, and a review by TRIVDA before a collection's first public listing. See Data model.
3. Earn
- You earn when another organisation's paid search returns passages from your dataset. Searching your own data earns nothing, and free searches pay no royalty.
- Your share is 70% of what the search paid; TRIVDA keeps 30%. When a search returns passages from several datasets, the payment is split by each dataset's share of the tokens returned.
- Earnings are recorded in your organisation's ledger as they happen, each row stamped with the share in force at the time.
Follow them in the dashboard under Earnings, or with the API:
GET /api/v1/datasets/usage: searches, passages served and earnings per dataset.GET /api/v1/datasets/{dataset_id}/earnings: one dataset's earnings per day and per buyer segment (industry, kind, country and size, at least three organisations per group; buyers are never named).GET /api/v1/usage/ledger: money earned and spent per day.GET /api/v1/usage/wallet: your earnings balance next to your spendable credits.
Earnings are paid out; they cannot be spent on searches. You may move them into credits with POST /api/v1/wallet/convert-earnings, after which they can no longer be paid out.
4. Monthly statements
After each month TRIVDA prepares a statement for every partner with earnings, from the ledger. It lists, per dataset, the gross amount, TRIVDA's share and your net earnings, and adds the VAT treatment for your invoice.
- Hold. A month's earnings are held for 30 days after the month ends, so refunds and chargebacks can be reversed first.
- Minimum. Below €50, the amount is carried over into the next month's statement instead.
- Issued. The statement gets a number and is frozen. You find it under Earnings → Statements (also as CSV), or with
GET /api/v1/statements.
Statements need your organisation's billing details (company name, address, country, invoice e-mail, and for Swedish companies the organisation number; an EU VAT number is checked in VIES). Add them under Settings → Billing or with PUT /api/v1/billing/profile.
5. Invoice TRIVDA
TRIVDA does not self-bill. When a statement is issued, you send TRIVDA an invoice for its amount:
- Invoice the statement's amount in EUR, and quote the statement number. TRIVDA's company details and the address for invoices are printed on the statement.
- Apply the VAT treatment the statement states for your organisation:
- Sweden: add Swedish VAT at 25%.
- Another EU country, with a VAT number confirmed in VIES: no VAT; write “Reverse charge” and both VAT numbers on the invoice.
- Outside the EU: no VAT; the invoice is outside the scope of EU VAT.
- TRIVDA records your invoice, pays it by bank transfer and marks the statement paid. The payout is written to your ledger, and your earnings balance goes down by the same amount.
Each statement has a digest, and its total equals the sum of your ledger rows for that month, so it can be checked against the ledger at any time.
During the pilot, partners are companies. Statements and payouts are manual for the pilot; automatic payouts follow before the public launch, and you will still invoice TRIVDA.