Privacy Policy
Last updated September 13, 2026
Token Delivery Network (“Token Delivery,” “we,” “us”) operates a hosted inference service that serves open-weight language models through an API. This Privacy Policy explains what data we collect when you use tokendelivery.ai and our API, how we use it, and the choices you have.
1. Scope
This policy applies to:
- Our websites (tokendelivery.ai, tokens.delivery, and their subdomains), including the playground.
- Our inference API at api.tokendelivery.ai.
- Any account interactions you have with us directly or through an aggregator or reseller.
It does not cover third-party services that connect to Token Delivery (such as aggregators, agent frameworks, or IDE extensions); those services have their own privacy policies.
2. Data We Collect
2.1 Account data. When you sign up we receive, through our sign-in provider, your name, email address, profile image, and the sign-in method you used. We do not see or store your password. Payments are processed by Stripe, which stores your card details; we keep only Stripe's customer and payment-method identifiers, never card numbers.
2.2 API keys. Keys are shown to you once and stored only as a cryptographic hash, with the key’s name, creation date, and rate limit.
2.3 API request data. When you call our API, your prompts, completions, images, video, and tool-call payloads are processed in memory to serve the request. They are not written to durable storage (see Section 5). The metadata we keep about each request is: the model, timestamps, token counts (prompt, completion, and cached), latency, the request id, the API key id, the region and machine that served it, and the HTTP status. At the network edge, your IP address and user agent are used for security and rate limiting.
2.4 Contact and feedback. Anything you send through the Contact window is stored with your email address if you are signed in, so that a person can read and answer it.
2.5 Website data. The websites use no advertising or analytics trackers. Cookies are used only for sign-in sessions and to remember your preferences on the site. Our edge and hosting providers keep standard, short-lived access logs.
3. How We Use Data
We use the data above to:
- Serve your inference requests and return responses.
- Operate, monitor, and improve the reliability, performance, and security of our systems.
- Investigate abuse, fraud, and violations of our Terms of Service.
- Meter usage, and bill for it once billing starts.
- Communicate with you about your account, service changes, and (if you opt in) product updates.
4. No Training on Your Data
We do not train, fine-tune, evaluate, or otherwise improve any model on your prompts or completions, and we do not share them with any model provider. This includes the small draft models we use for speculative decoding: they are trained on public datasets, never on customer traffic. Speculative decoding only accelerates generation; every token you receive is produced by the model you selected.
We use aggregated, non-content metadata (token counts, latency distributions, error rates) to plan capacity and improve our serving infrastructure.
5. Data Retention
Prompts and completions are not written to durable storage. While a request is being served they exist in the memory of the machine serving it. To speed up follow-up requests, that machine may keep derived attention state for a prompt (the model’s internal cache, not your text) in GPU or host memory for a short period; it is discarded when evicted, typically within hours, and whenever the machine restarts. It is never written to disk.
Request metadata (Section 2.3) is retained for billing reconciliation, capacity planning, security, and abuse prevention.
Account data is retained for as long as your account is active and, after closure, for the period required by applicable tax and accounting laws. Feedback is retained until it has been handled.
6. Zero Data Retention
Zero data retention is the default for every request to the Service, for every model, with no flag or special key required: prompts and completions are not stored after the request is processed, and only the operational metadata in Section 2.3 is kept.
7. Sharing and Disclosure
We share data only as needed to run the Service:
- Infrastructure providers. The API gateway, edge network, and artifact storage run on Cloudflare; the websites on Vercel; sign-in on Clerk; account and usage records on a managed PostgreSQL database; API-key and rate-limit state on a managed Redis service. Inference runs on GPU machines we operate on hosting providers and GPU marketplaces; your request data reaches those machines only in memory, over connections signed by our gateway, and those machines accept no other traffic.
- Aggregators and resellers. If you reach us through an aggregator such as OpenRouter, it receives the responses to your requests and handles your account and billing under its own policy.
- Legal requests. We may disclose data to comply with valid legal process, to protect our rights or property, or to protect the safety of users or the public.
- Business transfers. If Token Delivery is involved in a merger, acquisition, or asset sale, your data may be transferred as part of that transaction, subject to this policy.
We do not sell personal data, and we do not share prompts or completions with advertisers or data brokers.
8. Security
We use TLS for all data in transit, store API keys only as hashes, sign every request between our gateway and the machines that serve it, and restrict those machines to gateway traffic. No system is perfectly secure; we cannot guarantee absolute security, but we work to meet or exceed industry norms for inference providers.
9. Your Rights and Choices
Depending on where you live, you may have rights to access, correct, delete, or port your personal data, or to object to or restrict certain processing. To exercise these rights, email us at the address below. We will verify your identity before acting on a request.
You can also:
- Delete your API keys at any time from your account.
- Request deletion of your account and associated records, subject to legal retention requirements.
- Opt out of product emails using the unsubscribe link in any such email.
10. International Users
Inference is served from GPU machines in the United States and Canada, and occasionally from other regions where we rent capacity. Account and usage records are stored in the United States. If you access the Service from elsewhere, your data will be transferred to and processed in those locations, which may have different data protection laws than your jurisdiction. By using the Service, you consent to this transfer.
11. Children’s Privacy
The Service is not directed to children under 13 (or the equivalent minimum age in your jurisdiction), and we do not knowingly collect personal data from them. If you believe a child has provided us personal data, contact us and we will delete it.
12. Changes to This Policy
We may update this policy from time to time. We will post the new version at tokendelivery.ai/legal/privacy and update the “Last updated” date. For material changes, we will provide additional notice (for example by email or an in-product notice) before the changes take effect.
13. Contact Us
Questions, requests, or complaints about this policy can be sent to:
- Token Delivery Network
- Email: legal@tokendelivery.ai