Privacy Policy

Last updated September 13, 2026

Token Delivery Network (“Token Delivery,” “we,” “us”) operates a hosted inference service that serves open-weight language models through an API. This Privacy Policy explains what data we collect when you use tokendelivery.ai and our API, how we use it, and the choices you have.

1. Scope

This policy applies to:

It does not cover third-party services that connect to Token Delivery (such as aggregators, agent frameworks, or IDE extensions); those services have their own privacy policies.

2. Data We Collect

2.1 Account data. When you sign up we receive, through our sign-in provider, your name, email address, profile image, and the sign-in method you used. We do not see or store your password. Payments are processed by Stripe, which stores your card details; we keep only Stripe's customer and payment-method identifiers, never card numbers.

2.2 API keys. Keys are shown to you once and stored only as a cryptographic hash, with the key’s name, creation date, and rate limit.

2.3 API request data. When you call our API, your prompts, completions, images, video, and tool-call payloads are processed in memory to serve the request. They are not written to durable storage (see Section 5). The metadata we keep about each request is: the model, timestamps, token counts (prompt, completion, and cached), latency, the request id, the API key id, the region and machine that served it, and the HTTP status. At the network edge, your IP address and user agent are used for security and rate limiting.

2.4 Contact and feedback. Anything you send through the Contact window is stored with your email address if you are signed in, so that a person can read and answer it.

2.5 Website data. The websites use no advertising or analytics trackers. Cookies are used only for sign-in sessions and to remember your preferences on the site. Our edge and hosting providers keep standard, short-lived access logs.

3. How We Use Data

We use the data above to:

4. No Training on Your Data

We do not train, fine-tune, evaluate, or otherwise improve any model on your prompts or completions, and we do not share them with any model provider. This includes the small draft models we use for speculative decoding: they are trained on public datasets, never on customer traffic. Speculative decoding only accelerates generation; every token you receive is produced by the model you selected.

We use aggregated, non-content metadata (token counts, latency distributions, error rates) to plan capacity and improve our serving infrastructure.

5. Data Retention

Prompts and completions are not written to durable storage. While a request is being served they exist in the memory of the machine serving it. To speed up follow-up requests, that machine may keep derived attention state for a prompt (the model’s internal cache, not your text) in GPU or host memory for a short period; it is discarded when evicted, typically within hours, and whenever the machine restarts. It is never written to disk.

Request metadata (Section 2.3) is retained for billing reconciliation, capacity planning, security, and abuse prevention.

Account data is retained for as long as your account is active and, after closure, for the period required by applicable tax and accounting laws. Feedback is retained until it has been handled.

6. Zero Data Retention

Zero data retention is the default for every request to the Service, for every model, with no flag or special key required: prompts and completions are not stored after the request is processed, and only the operational metadata in Section 2.3 is kept.

7. Sharing and Disclosure

We share data only as needed to run the Service:

We do not sell personal data, and we do not share prompts or completions with advertisers or data brokers.

8. Security

We use TLS for all data in transit, store API keys only as hashes, sign every request between our gateway and the machines that serve it, and restrict those machines to gateway traffic. No system is perfectly secure; we cannot guarantee absolute security, but we work to meet or exceed industry norms for inference providers.

9. Your Rights and Choices

Depending on where you live, you may have rights to access, correct, delete, or port your personal data, or to object to or restrict certain processing. To exercise these rights, email us at the address below. We will verify your identity before acting on a request.

You can also:

10. International Users

Inference is served from GPU machines in the United States and Canada, and occasionally from other regions where we rent capacity. Account and usage records are stored in the United States. If you access the Service from elsewhere, your data will be transferred to and processed in those locations, which may have different data protection laws than your jurisdiction. By using the Service, you consent to this transfer.

11. Children’s Privacy

The Service is not directed to children under 13 (or the equivalent minimum age in your jurisdiction), and we do not knowingly collect personal data from them. If you believe a child has provided us personal data, contact us and we will delete it.

12. Changes to This Policy

We may update this policy from time to time. We will post the new version at tokendelivery.ai/legal/privacy and update the “Last updated” date. For material changes, we will provide additional notice (for example by email or an in-product notice) before the changes take effect.

13. Contact Us

Questions, requests, or complaints about this policy can be sent to: