manufact / MoteOS Privacy Policy
⚠️ This is a DRAFT. It must not be published as a binding policy, or used to obtain user consent, until it has been reviewed by counsel. Items marked
[TBC: …]are open placeholders and must all be resolved before launch. A consolidated list appears in Section 19.Drafting convention. This document describes what the product is expected to do on the date the policy takes effect. Where a collection is planned but not yet implemented in code, that is stated explicitly. Before launch, every such item must be re-verified; anything still unimplemented must be removed from this document rather than left standing as a description of behaviour that does not occur.
1. Scope
manufact is an AI workspace. You work with AI agents that run inside a sandbox dedicated to you, where they autonomously execute code, read and write your files, access the internet, and operate applications. MoteOS is the manufact desktop client (macOS and Windows); a browser version is also available.
This policy explains what information we collect, why, how we use and share it, how long we keep it, and what rights you have.
It applies to:
- the MoteOS desktop client (including the companion browser component shipped with it) and the browser client;
- the manufact backend services and your dedicated sandbox environment;
- the manufact / MoteOS product website and landing pages (the waitlist sign-up page, for example).
It does not apply to:
- third-party services you connect yourself through connectors (GitHub, GitLab, Gitee, databases, remote MCP servers, and so on) — those are governed by their own privacy policies;
- external websites and APIs that you or your agent choose to visit from the sandbox;
- content published by third parties that you install from the app, theme, or MCP marketplaces.
How this policy relates to the corporate website policy
This policy covers the MoteOS / manufact product and services. Our corporate website at https://www.archaea-ai.com has its own separate privacy policy (https://www.archaea-ai.com/privacy, effective 2026-08-07) that applies only to that website — web form submissions and website analytics. The two policies coexist and are independent: what happens inside the product is governed by this policy, and your visits to the corporate website are governed by the website policy. On anything product-related, this policy prevails.
This is an invite-only pre-release (beta). Please read Section 16, which sets out what that means for your data.
2. Who we are and how to reach us
| Data controller | Archaea AI, Inc. |
| Registered and business address | 149 Commonwealth Dr, Ste 1090, Menlo Park, CA 94025, USA |
| State of incorporation and company number | [TBC: not published on the corporate website; to be supplied by the business] |
| Contact address | sales@archaea-ai.com |
| Data Protection Officer | [TBC: whether a DPO is appointed, and contact details if so] |
| EU representative (GDPR Art. 27) | [TBC: we have no establishment in the EU; if the service is offered to EEA users an Art. 27 representative is normally required — counsel to confirm and complete] |
| UK representative (UK GDPR) | [TBC: as above, counsel to confirm whether a UK representative is required] |
You can contact us at the address above about anything in this policy, including the rights described in Section 11. It is the only contact address we currently publish; the operational task of setting up dedicated privacy, legal, and abuse aliases before launch is recorded in Section 19.
3. What we collect
We collect only what we need to run the service, keep it secure, and comply with the law. The table below is an overview; Sections 3.1 to 3.9 give the detail.
| # | Category | Main fields | Source | Primary purpose | Legal basis (GDPR Art. 6) |
|---|---|---|---|---|---|
| 1 | Account and identity | Subject identifier, email address, name and avatar URL returned by your identity provider; or the email address you use for a sign-in link | You / your identity provider | Account creation, sign-in, account linking | Contract 6(1)(b) |
| 2 | Access control | Waitlist email, referral source and UTM parameters, invite code, invite redemption record | You / generated by us | Managing rollout, preventing abuse | Contract 6(1)(b); legitimate interests 6(1)(f) |
| 3 | Network and approximate location | IP address, country/region inferred from IP, User-Agent | Network requests / CDN edge | Security, rate limiting, abuse prevention, regional statistics | Legitimate interests 6(1)(f) |
| 4 | Device and installation identifiers | Install id, machine id, MAC address, hostname, operating system and version, CPU architecture, client version, platform (desktop/browser) | The client | Detecting duplicate sign-ups and invite-code abuse | Legitimate interests 6(1)(f) — see Section 6 |
| 5 | Usage and billing | Input, output, cache-write and cache-read token counts; credits consumed; model name; conversation and message identifiers; timestamps | Generated by the system | Metering, quota enforcement, billing, capacity planning | Contract 6(1)(b) |
| 6 | Your content | Conversations, uploaded and generated files, sandbox files, microphone audio, images and video, text you select for the selection toolbar | You / your agent | Delivering the service itself | Contract 6(1)(b); see also 3.9 |
| 7 | Logs and audit records | Sign-in / sign-out records (time, event type, success and failure reason, IP, User-Agent, request id); system operation audit; server logs | Generated by the system | Security, troubleshooting, audit | Legitimate interests 6(1)(f); legal obligation 6(1)(c) |
| 8 | Local client data | Session token, interface preferences, local caches, in-session clipboard history | Stored on your device | Keeping you signed in, restoring your workspace | Contract 6(1)(b) |
We do not collect:
- passwords — manufact is passwordless; there is no sign-up password and we store no password or password hash;
- card numbers or bank account details — the beta is free of charge;
[TBC: once paid plans launch, payment details are handled by the payment provider and we do not see full card numbers]; - precise location (GPS or street-level address);
- camera imagery — the main client requests no camera permission and contains no code path that opens a camera;
- screen recordings or screenshots of your desktop — the client requests no Screen Recording permission and contains no code path that captures your screen;
- contacts, calendars, photo libraries, installed-application lists, or browsing history;
- political opinions, religious beliefs, health data, biometric templates or other special category data — unless you choose to upload such material as content (see 3.6).
3.1 Account and identity
manufact uses no passwords. You sign in one of two ways:
- Third-party single sign-on: Google, Microsoft, or Apple. From your identity provider we receive your subject identifier within that provider, your email address, your name, and an avatar URL where the provider returns one. We never receive or store your password with that provider.
- Email sign-in link (magic link): you give us an email address and we send a one-time sign-in link to it. We store that address.
Where two sign-in methods resolve to the same verified email address, they are automatically linked to a single account (one account may hold several identity records).
Your session token is an opaque random string (not a parseable JWT), held server-side with a limited lifetime and sliding expiry.
3.2 Access control (waitlist and invite codes)
The beta uses a waitlist and invite codes to manage rollout:
- Waitlist: email address, status, referral source and UTM parameters, a hash of the IP address and the User-Agent at submission, a hash of the confirmation token, and timestamps for each stage.
- Invite codes: the code itself, batch label, the email it was issued to (if any), maximum and used redemption counts, the plan and bonus credits it grants, and its expiry.
- Redemption records: the account email, sign-in method (google / microsoft / apple / email), platform (desktop or browser), IP address and IP hash, IP country/region, client version, device fingerprint and device descriptor, User-Agent, and the redemption timestamp.
Two things stated plainly
- By design, the IP address in a redemption record is stored both in clear text and as a hash. The hash is what lets us correlate against other hashed logs; the clear text is what an actual abuse investigation needs. The device descriptor is stored in clear text (see 6.1). We disclose this rather than describing the record as “anonymised”.
[TBC: as of this draft (2026-08-25) the database columns for these forensic fields exist, but neither the client-side reporting nor the server-side write logic is implemented, and nothing computes the IP hash. Before launch, verify what is actually collected and delete any field from this section that is not.]
3.3 Network and approximate location
- We record the IP address a request comes from. Where traffic passes through a content delivery network, we read the originating client IP from the header the edge provides.
- We record the ISO country/region code the CDN edge derives from that IP. We do not collect GPS coordinates, precise location, or street-level addresses.
[TBC: the country-code lookup is not yet implemented; verify before launch] - We record the User-Agent string sent by your browser or client.
- These are used for rate limiting and denial-of-service protection, detecting unusual sign-ins and bulk registration, and making regional compliance and capacity decisions.
3.4 Device and installation identifiers
When you register or redeem an invite code, the desktop client generates and reports a set of device identifiers used to detect repeat sign-ups from the same machine. These identifiers are stored in clear text (with a separate hash of the IP address kept for correlating logs) — a deliberate choice, whose reasons, limits, and your right to object are set out in Section 6. That is the part of this policy we most want you to read.
3.5 Usage and billing
After each agent run we record input tokens, output tokens, cache-write and cache-read tokens, the model used, the credits computed from them, the associated conversation and message identifiers, the raw usage object, and a timestamp.
These records drive metering and quota enforcement, billing reconciliation, and capacity planning. They record measurements of usage; they do not contain the text of your conversations.
3.6 Your content
“Your content” means:
- what you type into a conversation, images you paste, and files you upload;
- files your agent creates, modifies, or reads inside your sandbox;
- microphone audio captured by the voice features (dictation, voice sessions, live calls);
- prompts and reference material you submit to image or video generation, and the results;
- data on third-party services that you authorise us to access on your behalf through a connector (repository contents, query results, and so on);
- documents you upload for parsing (for example, PDF to Markdown conversion).
Your content is transmitted to third-party model and speech providers for processing. That is an unavoidable technical premise of an AI product; see Section 5.
About voice recordings. To avoid repeated provider calls and to support playback, the server caches your recorded audio clips and synthesised speech on our own server disks, expiring them automatically after a set period. These caches do not go to object storage or a public CDN and are used only for your own sessions. [TBC: retention period for voice and TTS caches]
We do not scan your sandbox files for any purpose other than delivering the service. The one exception is content you choose to submit for public publication to a marketplace, which is reviewed automatically (see 8.4).
If you upload sensitive personal information about other people — identity documents, medical records, biometric data — that is your choice. Please first read the restriction on strictly regulated data in the Disclaimer and Terms of Use.
3.7 Logs and audit records
- Sign-in and sign-out audit log: user id and username, event type, success or failure and the failure reason, sign-out reason, IP address in clear, User-Agent in clear, request id, and timestamp. These fields are not hashed. The log is readable by administrators only.
- System operation audit: user id, session and container ids, action, resource type and id, structured detail, timestamp.
- Server logs: request paths, status codes, latency, and error traces, which may contain IP addresses and request ids.
3.8 Data stored locally on your device
The desktop client keeps some data on your own computer. Unless this policy says otherwise, none of it is uploaded.
| Where | What |
|---|---|
| System keychain (macOS Keychain / Windows Credential Manager) | One item only: your backend session token |
Browser local storage (localStorage, namespaced per account) |
Interface preferences (theme, accent, language, font size), window and desktop layout, Dock and file-browser preferences, recently opened file paths, installed theme packs and wallpaper references, avatar cache, backend address, update bookkeeping, and file-sync pairing details including the absolute path of the local folder you chose |
| Application data directory | Wallpaper copies, font cache, logs, file-sync component state |
| Memory only (never written to disk) | In-session clipboard history: up to 50 entries from copy and cut actions performed inside MoteOS, held in memory only, never written to disk and never uploaded, cleared when you sign out. It does not record copies you make in other applications. |
3.9 Operating system permissions and on-device capabilities
The client asks your operating system for the permissions below. Each takes effect only once you grant it, and each serves one specific feature.
| Permission | When we ask | Purpose and limits |
|---|---|---|
| Microphone | The first time you use a voice feature | Voice dictation, voice sessions, and live calls. Audio only — no video. Audio is sent to a speech provider (see 5.1) and cached server-side for a limited period (see 3.6). |
| Local network access | When connecting to a backend on your local network | Connecting to a manufact backend you or your organisation runs, and syncing files with devices on your network. |
| Accessibility (macOS) | Only if you turn on the desktop selection toolbar in Settings | See below. This feature is off by default. |
About the desktop selection toolbar (important)
This is the only feature in the client that reads content from other applications, so it warrants its own explanation:
- It is off by default. It runs only after you enable it in Settings and grant Accessibility permission; without that permission it refuses to start.
- How it triggers: the client listens only for left mouse button press and release, which it uses to recognise a double-click-to-select or a drag-select. It does not monitor the keyboard and does not record keystrokes.
- What it reads: after you complete a selection gesture, it reads the text you selected in the front-most application, along with that application’s name and bundle identifier. If the system accessibility interface cannot return the selection directly, the client falls back to triggering that application’s own Copy menu item, reads the clipboard, and then restores the clipboard’s previous contents in full.
- Where the text goes: the text is shown only in a floating toolbar on your screen. It is sent to our servers only when you click an action (translate, summarise, explain, search, read aloud); “copy” is handled entirely locally. Diagnostic records store the application identifier and timing, never the text itself.
- Exclusion list: you can exclude any application, and excluded applications are never read.
[TBC: the exclusion list currently ships empty. We strongly recommend pre-populating it with password managers, terminals, and banking applications before launch, and showing a clear explanation the first time the feature is enabled.] - A known side effect: to wake the accessibility tree in some applications (Electron and Chromium-based ones), the client sets accessibility-related flags on those processes and does not clear them afterwards. This reads no additional content, but it leaves that application’s accessibility interface active for the rest of its run.
About the clipboard
The client reads your system clipboard only within an action you explicitly trigger: pressing paste or choosing Paste from the menu, copy and paste inside the terminal, and the read-and-restore sequence in the selection feature described above. There is no clipboard polling and no background clipboard monitoring of any kind.
About files on your computer
The client does not scan your Documents, Desktop, or Downloads folders, does not enumerate your installed local fonts, and does not read any file you have not chosen. The only ways local files are touched are: files you pick through a system dialog (uploads, wallpaper import, font import), files you drag into a window, and the folder you designate for file sync. If you enable local tool acceleration (off by default), the client runs a bundled media-processing binary on your machine, with inputs downloaded from and outputs uploaded back to our servers.
About automatic updates
At startup and every six hours thereafter, the client requests a static version manifest file from our update server. The request carries no version number, device identifier, account information, or parameters of any kind — the version comparison happens locally on your device. Like any HTTP server, the update server can observe your IP address, User-Agent, and request timing. Update packages are verified against a digital signature before installation.
About the companion browser shipped with the client
The desktop client bundles a separate browser component for agent use. [TBC: the component's application manifest currently carries default camera and Bluetooth permission declarations generated by the packaging tool, which the feature does not use. These unused declarations should be removed before launch; if they are kept, this policy must disclose that macOS may present the corresponding permission prompts.] The component defaults to Google search and can install extensions from the Chrome Web Store; using those features sends the corresponding requests directly to Google.
About file sync (optional)
If you enable file sync, the client runs a bundled sync component that keeps a local folder you choose in step with your file volume on the server. Local network broadcast discovery is disabled in the configuration we generate. [TBC: the component's global discovery servers and relay pools are not explicitly disabled. Left at their defaults, your device would communicate with that open-source project's public discovery and relay servers. These should be explicitly disabled before launch, or disclosed here.]
What we have not integrated. Neither the client, the backend, nor the product website contains any third-party advertising SDK, cross-site tracking pixel, product analytics SDK, or crash-reporting SDK. This has been verified field by field — no Sentry, PostHog, Mixpanel, Google Analytics, Umami, Plausible, or equivalent. (The corporate website at archaea-ai.com does use Google Analytics; that is governed by its own privacy policy — see Section 1.) [TBC: if any such service is added before launch, it must be disclosed here with the provider, the fields collected, and how to opt out.]
4. How we use this information
| Purpose | Categories | Legal basis |
|---|---|---|
| Creating and maintaining your account; keeping you signed in | 1, 8 | Contract |
| Providing agents, sandboxes, files, voice, and generation features | 6, 5 | Contract |
| Metering usage, enforcing plan quotas, billing | 5, 1 | Contract |
| Managing beta rollout; validating and redeeming invite codes | 2, 3, 4 | Contract; legitimate interests (abuse prevention) |
| Preventing duplicate sign-ups, free-quota farming, and other abuse | 3, 4, 7 | Legitimate interests — see Section 6 |
| Securing the service; detecting and responding to attack and fraud | 3, 4, 7 | Legitimate interests; legal obligation |
| Diagnosing faults, restoring service, improving reliability | 5, 7, 4 | Legitimate interests |
| Contacting you about service changes, security incidents, account status | 1 | Contract; legitimate interests |
| Sending product news and marketing email | 1 | Consent (unsubscribe at any time) |
| Meeting legal obligations and responding to lawful requests | All | Legal obligation |
We do not carry out solely automated decision-making or profiling that produces legal or similarly significant effects. Anti-abuse rules can lead to an account being suspended, but suspension and appeal both involve human review (Section 11).
5. How your content is processed
5.1 Content goes to third-party providers
For agents to work, we must send the relevant content to third parties that provide model capabilities:
| Situation | What is transmitted | Recipient category |
|---|---|---|
| Conversations and agent runs | Your messages, system prompts, the contents of files the agent reads, tool calls and results | Large language model providers |
| Conversation title generation, marketplace review | The relevant text | Large language model providers |
| Speech-to-text | Your microphone audio (as a file or a stream) | Speech recognition providers |
| Text-to-speech | The text to be spoken | Speech synthesis providers |
| Live voice calls | Two-way audio streams and recognised transcripts | Realtime voice providers |
| Image and video generation | Prompts and reference images; results are downloaded from the provider | Image and video generation providers |
| Selection translate / summarise / explain | The text you selected (only after you click an action) | Large language model providers |
| Submitting content for public publication | The submitted text and images | Model providers; content-safety providers |
| Document parsing (for example, PDF to Markdown) | The document you uploaded | Our own document parsing service |
Note in particular: when an agent reads a file from your sandbox in order to carry out your instruction, the contents of that file become part of the request sent to the model provider. This is what makes it possible for an agent to work on your files at all.
The providers we use: [TBC: the provider list to be disclosed at launch. We recommend disclosing by category plus a link to a publicly maintained sub-processor list, so that changing a supplier does not require re-obtaining consent. It must match our actual contracts with those providers.]
5.2 Bring your own key (BYOK)
Some plans let you configure your own model API keys. Two differences matter:
- Agent main-model key: a key you configure for the agent’s main model is injected as an environment variable into your sandbox container, so the agent inside the container can use it directly.
- Keys for other purposes (utility models, image, video): stored server-side encrypted with AES-256-GCM and never placed in the sandbox.
In either case we do not use your key for anything other than your own requests. Whether the upstream provider uses your content for training is governed by your agreement with that provider, which we cannot control on your behalf.
5.3 Whether your content is used for training or model improvement
We do not use your conversations, files, or audio to train or fine-tune any model, and we do not provide them to third parties for training. We have disabled training-related options with our upstream providers, or contracted with them not to use our data for model training.
If we ever want to use content for model improvement, we will ask for your explicit consent first, provide a switch you can turn off at any time, and will not begin before you agree.
The one exception is bring-your-own-key (BYOK). Where you use your own model API key, whether that upstream provider uses your content for training is governed by your agreement with that provider, which we cannot control on your behalf (see 5.2).
5.4 Access by our people
Our staff and contractors have no default access to your conversations or sandbox files. The exceptions, each of which is logged:
- you explicitly authorise us to look at specific content while helping you with a support request;
- access is necessary to diagnose a severe fault affecting service availability, minimised as far as possible;
- the law requires it, or it is necessary to investigate serious abuse of the terms of use.
6. Device identifiers and abuse prevention
This gets its own section because device identification is the most intrusive collection described in this policy, and you are entitled to know exactly where its boundaries lie.
6.1 What we collect
When you register or redeem an invite code, the desktop client generates and reports a device descriptor which may include:
- an install id — a random identifier generated the first time the client runs, stored in the system keychain or local storage;
- a machine id — a machine-level identifier provided by the operating system;
- the MAC address of a network interface;
- the hostname;
- the operating system type and version, and CPU architecture;
- the client version.
The server stores one stable device identifier in an indexed field for comparison, and the reported descriptor as a whole in a structured field.
Two things stated plainly
- Raw device identifiers are stored in clear text. What the server stores is the raw descriptor object reported by the client: the MAC address, hostname, machine id, operating system and version, client version, and install id are all written to the database in clear text, with no hashing or redaction. We separately keep a hash of the IP address for correlating records across logs. This is a deliberate choice, not an oversight — we need human-readable raw values for forensic review and for handling appeals. Why we do not store a digest instead, and what limits we place on this data, are set out in 6.4(c) and (d).
[TBC: as of this draft (2026-08-25), the client does not implement device identifier collection or reporting and the server has no write logic — the database columns exist and the libraries are vendored, but no code computes, transmits, or stores any device identifier. Before launch: if this is still unimplemented, delete this entire section; if the implemented field set differs from the list above, rewrite it to match.]
6.2 Why we collect it
The sole purpose is detecting duplicate sign-ups from the same machine and invite-code abuse.
During the beta every account receives free credits, and those credits correspond to real costs we pay upstream model providers. In a passwordless model, “registering” is simply “signing in for the first time”, and email addresses can be created without limit, including disposable ones. Email alone cannot identify one person opening accounts repeatedly. IP alone cannot either — shared exits on home, corporate, and campus networks would treat unrelated people as the same person, causing more harm than it prevents. In this specific situation, a device identifier is the available measure with the lowest false-positive rate.
6.3 What we will not use it for
- Not for advertising, ad attribution, or cross-site or cross-app behavioural tracking of any kind;
- Not sold, rented, or provided in any form to data brokers or ad networks;
- Not used to build an interest profile or infer personal characteristics about you;
- Not used to identify you on other websites or applications;
- Not collected at any point other than registration — no periodic reporting, no continuous tracking.
6.4 Legitimate interests balancing (GDPR Art. 6(1)(f))
We rely on legitimate interests for this processing and have carried out the following assessment.
(a) The interest is legitimate. Preventing bulk farming of free quota bears directly on whether the service can continue to exist and whether ordinary users get the quota and response times they are entitled to. Abuse and fraud prevention is a widely recognised legitimate interest of a controller, and GDPR Recital 47 names fraud prevention explicitly.
(b) Necessity and minimisation.
- Collected once, at a single moment — registration or invite redemption — and never reported again;
- Limited to fields needed to establish that two sign-ups came from one machine; we do not collect contacts, installed application lists, or browsing history;
- We do not use covert browser-fingerprinting techniques such as canvas fingerprinting or font enumeration;
- We assessed weaker alternatives (email verification alone, IP alone, CAPTCHA alone) and concluded that each either fails to achieve the purpose or causes disproportionate false positives.
(c) Why we store this in clear text rather than as a digest.
This is the point we most owe you an explanation for. We deliberately store raw device identifiers — MAC address, hostname, and the rest — in clear text rather than keeping only a hashed digest of them. Our reasons, which we would rather set out honestly than paper over with a comfortable line about “storing only a digest”:
- The actual use of this data is forensic review by a person, not automated banning. A digest can only answer “are these two records byte-for-byte identical?” It cannot tell a reviewer which field matched and which did not — for instance, that the MAC address changed after a network card was replaced while the machine id and hostname stayed the same — so it cannot support the human judgement between “the same machine registering twice” and “two similarly configured machines”.
- Appeals require explainability. When a user challenges a restriction imposed for abuse prevention, we must be able to explain what the finding rests on and let them rebut it. With only a digest, all we could produce is an unreadable hash. Taking explainability out of the process hurts the wrongly flagged user most, not the abuser.
- A digest adds little real privacy here. Device identifiers come from a small, structured value space, so anyone holding the stored data can recover the original by brute force. “Digest only” would be largely a formal protection rather than a substantive one.
(d) The minimisation and limits that go with it.
- A single purpose: detecting duplicate sign-ups and invite-code abuse only. Never advertising, ad attribution, cross-site or cross-app tracking, profiling, or inference about you;
- A single moment: collected once at registration or invite redemption — no periodic reporting, no continuous tracking;
- Never sold or passed on: not sold, rented, or provided in any form to data brokers or ad networks, and not used for any business purpose beyond our own abuse prevention;
- Restricted access: these records are readable by administrators only; internal access follows least privilege and is logged;
- Deleted with the account: when an account is deleted, the device identifier records associated with it are deleted too (see Sections 10 and 11);
- A defined retention period, after which the data is deleted or irreversibly anonymised (the period itself is in Section 10 and is still to be settled);
- Specific disclosure: every field and its storage form is listed in this policy, rather than hidden behind a vague reference to “device information”.
(e) Impact on you, and our conclusion.
- The identifiers are not used for advertising or cross-service tracking and are not shared, so this processing has no effect on your life outside manufact;
- The principal adverse impact is that these identifiers can link accounts to one another, and that a MAC address and hostname held in clear text are strongly identifying — we are not minimising that;
- Subject to the limits in (d), we consider that our legitimate interest is not overridden by your rights and freedoms. If an approach emerges that supports human forensic review and appeals while being less identifying, we will reassess and revise this section accordingly.
(f) Your right to object. You may object at any time to processing based on legitimate interests under GDPR Art. 21: write to the address in Section 2 with your account and what you are asking for. On receiving an objection we will delete the device identifier records we hold for you and stop collecting them; you may also ask to access or delete those records separately under Section 11. Please note: if you object to device identifier processing, we may be unable to offer you an account with free credits, because that allocation depends on this anti-abuse measure. It does not affect your use of a paid account.
7. Data processing arising from agent autonomy
manufact’s agents do not merely answer questions; they act on your behalf. That produces several processing facts specific to this product:
- Agents read and write files in your sandbox, including creating, modifying, and deleting them. Files they read are sent to model providers as part of the request (see 5.1).
- Agents access the internet: fetching pages, calling APIs, downloading packages, running searches. These requests originate from your sandbox, so the destination sees our server’s egress IP rather than your home IP — but it sees the content of the request. We have no control over how those third parties handle what you send them.
- Agents may operate browsers and application interfaces, and therefore process whatever is displayed there.
- Connectors (GitHub, GitLab, Gitee, databases, remote MCP servers, and others): once authorised, an agent can read — and potentially write — your data on that service within the scope granted. Credentials are stored encrypted with AES-256-GCM, and you can revoke them at any time in Settings.
- Marketplace browsing: when you browse the MCP marketplace, the client requests catalogue data from public MCP registries through our backend.
- Scheduled tasks: if you configure them, agents run to schedule while you are away, with all the processing that entails.
What you ask an agent to do, and what you connect it to, therefore determines in substance what data gets processed. Please consider that before authorising.
8. How we share information
We do not sell personal information and do not share it for targeted advertising. We share in the following circumstances.
8.1 Service providers (processors)
| Category | Purpose | Data involved |
|---|---|---|
| Large language model providers | Agent and text capabilities | Your content (category 6) |
| Speech providers (STT / TTS / realtime) | Voice features | Audio, text to be spoken, transcripts |
| Image and video generation providers | Generation features | Prompts, reference images, results |
| Identity providers (Google / Microsoft / Apple) | Single sign-on | The sign-in request; we receive identifier, email, name |
| Transactional email provider | Sign-in links, invitations, notifications | Email address and message content |
| CDN / WAF / DDoS protection | Security and delivery | IP, request metadata |
| Cloud infrastructure and hosting | Running the service and databases | All categories, as the substrate |
| Content-safety providers | Reviewing publicly published marketplace content | Submitted images and text |
| Payment providers | [TBC: once paid plans launch] |
[TBC] |
Provider list: [TBC: we recommend maintaining a publicly accessible sub-processor page and linking it here, with advance notice of changes.]
[TBC: Japan is served through an authorised distributor, クリューグル株式会社 (Krugle Inc.). If the distributor handles user personal information in the course of sales, invoicing, or local support, its role (independent controller or processor) and the data categories involved must be stated in the table above and an appropriate agreement signed. If it handles none, delete this note.]
We enter into data processing agreements with our processors, requiring them to process only on our instructions, keep the data confidential, and apply appropriate security measures. [TBC: before launch, confirm one by one that a DPA is in place with every processor (or that their standard DPA has been accepted). If any is outstanding, this sentence must not stand as an unqualified commitment.]
8.2 Legal requirements
Where we form a good-faith belief that disclosure is necessary to comply with the law, a court order, or a valid government request, we may disclose the relevant information. To the extent the law allows, we review such requests, require them to be in proper legal form, and try to notify you in advance.
8.3 Corporate transactions
In a merger, acquisition, or sale of assets, your information may transfer as part of that transaction. We will notify you beforehand and ensure the acquirer is bound by terms no less protective than this policy.
8.4 Content you publish yourself
If you publish or share a theme, app, or conversation to a marketplace or via a share link, it becomes visible to others within the scope you chose. Before publication such content may be reviewed automatically (static scanning, model-based review, and a third-party content-safety service).
Stated plainly: when a third-party content-safety service reviews images, the image is first uploaded to that provider’s own object storage before being scanned.
[TBC: whether third-party content-safety review is enabled during the beta — it is configurable and inactive when no keys are set. If not enabled, delete this paragraph.]
8.5 What we do not do
- We do not sell personal information to data brokers;
- We do not share personal information for third-party advertising;
- We do not provide your content to others for model training (see also 5.3).
9. International transfers
Our service is global, and our infrastructure and providers may sit outside your country, so your information may be transferred to other jurisdictions.
The controller, Archaea AI, Inc., is located in California, United States. Personal information of users in the European Economic Area, the United Kingdom, Switzerland, Japan, and Korea is therefore transferred to and processed in the United States.
- Server regions:
[TBC: production deployment regions] - Model and speech provider regions:
[TBC: to be settled with the provider list. If the list includes providers in mainland China, that requires separate treatment and explanation.]
For transfers of personal data out of the European Economic Area, the United Kingdom, or Switzerland, we rely on appropriate safeguards under Chapter V of the GDPR:
- an adequacy decision of the European Commission, where the destination has one;
- otherwise Standard Contractual Clauses
[TBC: modules used and execution date], together with any necessary supplementary measures (transfer impact assessment, encryption, access controls); - for the United Kingdom, the UK IDTA or the UK Addendum
[TBC].
You can request a copy of these safeguards at the address in Section 2; commercially confidential portions may be redacted.
10. Retention
We keep data no longer than is necessary for the purpose it was collected for.
| Data | Retention |
|---|---|
| Account and identity records | For the life of the account; after deletion [TBC: grace period, 30 days suggested] |
| Conversations and sandbox files | For the life of the account, or until you delete them [TBC: how long deleted content persists in backups] |
| Session and container records | [TBC: the backend default is currently 30 days; confirm the production value] |
| Voice recordings and TTS caches | [TBC: currently expire on a configurable TTL; confirm the production value] |
| Usage and billing records | [TBC: suggest the statutory tax and accounting period, typically 5–7 years; these records contain no conversation text] |
| Sign-in / sign-out audit logs | [TBC: 6–12 months suggested] |
| Server logs | [TBC: currently 30 days by default] |
| Invite redemption and device identifier records | [TBC: 12 months suggested; also state whether expiry deletes the plaintext IP and device descriptor while retaining the hash] |
| Waitlist records | [TBC: clearance rule once converted to an account, or on unsubscribe] |
After the retention period we delete or irreversibly anonymise the data. Copies held in backups expire with the ordinary backup rotation.
What account deletion carries with it: when your account is deleted, the invite redemption and device identifier records associated with it — including the clear-text MAC address, hostname, and the rest of the device descriptor — are deleted with it, rather than being kept for the periods in the table above.
11. Your rights
Wherever you are, we offer you the following rights:
- Access — find out what personal information we hold about you and obtain a copy;
- Rectification — correct information that is inaccurate or incomplete;
- Erasure — ask us to delete your personal information, subject to statutory retention obligations. Deleting your account also deletes the device identifier records described in Section 6, and you can ask us to delete those records on their own without closing your account;
- Restriction — ask us to pause processing in defined circumstances;
- Objection — object to processing based on legitimate interests, including the device identifier processing in Section 6. Objection to direct marketing is absolute;
- Portability — receive the data you provided in a structured, commonly used, machine-readable format, or have it transmitted to another controller;
- Withdrawal of consent — where processing rests on consent (marketing email, and any future model-improvement use), withdraw it at any time, without affecting the lawfulness of processing before withdrawal.
How to exercise them: contact us at sales@archaea-ai.com (the contact address we currently publish; setting up a dedicated privacy alias is an operational task recorded in Section 19). We respond within 30 days; where a request is complex we may extend that period and will tell you why within the initial 30 days. To protect your account we may need to verify your identity first. Exercising your rights is free, unless a request is manifestly unfounded or excessively repetitive.
[TBC: when self-service account deletion and data export will ship. As currently implemented there is no self-service route for either, so rights requests are handled manually by email. Before launch this should be stated on the help page, and the manual workflow confirmed to fit within 30 days.]
Complaints: if you believe our processing breaches applicable law, you may complain to the data protection authority where you live. [TBC: we have no establishment in the EU, so the one-stop-shop "lead supervisory authority" mechanism does not normally apply — counsel to confirm whether a GDPR Art. 27 representative is required and, on that basis, which authority users should be directed to] We would also like you to come to us first — it is usually faster.
California residents (CCPA / CPRA)
- You have the right to know the categories and specific pieces of personal information we collect, use, and disclose; to request deletion and correction; and not to be discriminated against for exercising those rights.
- We do not “sell” your personal information, and we do not “share” it for cross-context behavioural advertising as those terms are defined by the CCPA. We therefore do not offer a “Do Not Sell or Share My Personal Information” link, but you can still contact us about any of the above.
- We do not knowingly sell or share the personal information of anyone under 16.
Japan (APPI) and Korea (PIPA)
- We handle personal information of users in these regions in line with Japan’s Act on the Protection of Personal Information (APPI) and Korea’s Personal Information Protection Act (PIPA): telling you the purpose of use at the point of collection, meeting the applicable notice or consent obligations before providing data to a third party (including across borders), and handling the access, correction, deletion, and suspension-of-use requests listed in this section.
- Those requests go to the same address as everything else in this policy (Section 2).
Other jurisdictions: this policy is written to GDPR / UK GDPR, CCPA / CPRA, Japan’s APPI, and Korea’s PIPA — the same set of regimes our corporate website privacy policy addresses. We do not currently offer the service in jurisdictions such as mainland China (PIPL), Brazil (LGPD), or Canada (PIPEDA); if we do so in future, we will adapt this policy at that point.
12. Cookies and local storage
- Product website and browser client: we use strictly necessary cookies and local storage to keep you signed in and remember your interface preferences. The product website and landing pages are currently fully static, with no analytics or advertising scripts.
[TBC: if analytics cookies are introduced at launch, a consent banner is required and they must be listed here; if not, state plainly that we use no advertising cookies and no third-party analytics cookies.] - Corporate website (archaea-ai.com): that site uses Google Analytics. Its cookies and website analytics are governed by the website’s own privacy policy (see Section 1) and fall outside this policy.
- Desktop client: your session token lives in the operating system keychain and your interface preferences in local storage, namespaced per account. Neither is advertising technology, neither is used for cross-site tracking, and neither is readable by third parties.
- We serve no advertising, so there are no advertising cookies or advertising identifiers at all.
- About embedded third-party interfaces: the client embeds several web-based components (the online code editor, online document editing, and mini-apps you install). So that these can save your editing state properly, the client disables the browser engine’s default cross-site storage partitioning. The consequence is that these embedded components can persist their own local data on your computer. They are provided by us or by the publisher of what you installed and serve those components’ own functionality; the client loads no third-party advertising or analytics scripts.
- You can clear local storage in your browser, or sign out in the client to clear the token. Either will sign you out and reset your interface preferences.
13. Children
manufact is not intended for children, and we do not knowingly collect personal information from them. The age threshold depends on where you are:
- United States: consistent with the Children’s Online Privacy Protection Act (COPPA), the service may not be used by anyone under 13;
- European Economic Area and the United Kingdom: the digital age of consent set by your country applies (member states set it anywhere between 13 and 16). Below that age, a parent or guardian must consent and exercise the relevant rights on the child’s behalf;
- Elsewhere: the minimum age set by local law applies.
If you are a parent or guardian and believe your child has given us personal information without your consent, contact us at the address in Section 2; we will verify and then delete the information and close the account.
[TBC: counsel to confirm these thresholds against the final target markets — including whether the differing EEA digital ages of consent are handled by applying the lowest age uniformly or by country, and whether an age declaration or age verification step is needed in the sign-up flow.]
14. Security
The technical and organisational measures we take include:
- Kernel-level sandbox isolation: each user’s agents run in a dedicated gVisor sandbox container with a private file volume. Sandboxes and file volumes are not shared between users.
- Encryption in transit: communication between the client and the service we host uses TLS. (If you connect to a backend you run yourself, transport security depends on how that deployment is configured.)
- Encrypted credential storage: connector credentials and your own model keys are stored encrypted with AES-256-GCM.
- Passwordless authentication: we store no passwords, so there is no password database to leak. Session tokens are opaque random strings with a limited lifetime and sliding expiry.
- Access control and audit: internal access follows least privilege, and significant operations are logged.
- Network defences: layered rate limiting; edge WAF and DDoS protection
[TBC: to match the edge protection actually enabled in production]. - Privilege reduction: terminals and editors inside the sandbox run as a non-root user.
- Update integrity: client updates are verified against an Ed25519 (minisign) signature before installation.
We do not promise absolute security. No method of transmission over the internet or of electronic storage is completely secure, and we cannot guarantee absolute security. Please do your part too: protect your sign-in email and third-party accounts, and do not sign in on devices you do not trust.
Breach notification: in the event of a personal data breach likely to result in a high risk to you, we will notify the supervisory authority within the period the law requires (72 hours of becoming aware, under the GDPR) and notify you directly where the conditions for doing so are met.
15. Where your data lives
Your files, conversations, voice caches, and theme assets are stored on server disks we operate ourselves. We do not use third-party object storage or a public CDN to host your content. The exceptions are set out individually elsewhere in this policy (the model and speech providers in 5.1, and content-safety review in 8.4).
16. Beta-specific notice
This is an invite-only pre-release. Please be aware:
- Data may be lost. During the beta we rebuild environments, change data structures, and migrate infrastructure frequently. Your sandbox, files, and conversation history may be reset or lost, possibly irrecoverably.
- Keep your own backups. Do not make manufact the only place important material exists.
- There is no service level agreement. Features may change, be suspended, or be withdrawn at any time.
- Access can be revoked. We may withdraw invitations and access for abuse or capacity reasons.
- Logging during the beta may be more detailed than in general availability, so that we can diagnose problems. We will narrow it once the service is generally available.
- Some features are still evolving, and the scope of collection described here may change with them. Any change will be notified as set out in Section 17.
For the risks and liability boundaries of the service itself, see the Disclaimer and Terms of Use in the same directory.
17. Changes to this policy
- Ordinary revisions: we update the version number at the top of this document and publish the new text on the website and in the client before it takes effect.
- Material changes — new categories of collection, changes to purposes or third-party flows, changes to retention periods, or a change of legal basis — are notified at least 30 days in advance by email or a prominent in-client notice. Where a change depends on consent, we ask for your consent again and do not begin the processing before you agree.
- We keep an archive of previous versions, available on request at the address in Section 2.
18. Complaints and contact
For any question, comment, or complaint about this policy or our data practices:
- Email:
sales@archaea-ai.com - Postal address: Archaea AI, Inc., 149 Commonwealth Dr, Ste 1090, Menlo Park, CA 94025, USA
We aim to reply within [TBC: committed response time, 30 days suggested].
19. Open placeholders for the business and counsel
Consolidated list of every [TBC] above. This document must not go live until all of them are resolved.
Three items settled on 2026-08-25 and no longer open: (1) content is not used for training (§5.3, BYOK excepted); (2) raw device identifiers are stored in clear text — a deliberate choice, with the reasoning and limits in Section 6; (3) the legal entity is Archaea AI, Inc., with its address, contact details, and the set of applicable regimes aligned to the corporate website.
Entity and contact
- State of incorporation and company number (§2)
- Whether a DPO is appointed, and contact details (§2)
- Whether an EU (Art. 27) or UK representative is required, and contact details (§2)
- Which authority users should be directed to for complaints — we have no EU establishment, so the one-stop-shop lead supervisory authority does not normally apply (§11)
Effect and applicability
5. Effective date (frontmatter effective_date; the corporate website policy’s effective date of 2026-08-07 is available for reference)
6. Counsel to confirm the children’s age thresholds: whether the differing EEA digital ages of consent are applied uniformly at the lowest age or by country, and whether an age declaration or verification step is needed at sign-up (§13)
Third parties and data flows 7. The model / speech / image / video provider list to disclose at launch, and the form of disclosure (§5.1, §8.1) 8. Whether error monitoring, product analytics, or crash reporting is added; if so, provider, fields, and opt-out (§3.9, §8.1) 9. Payment providers and the scope of their processing (§3, §8.1 — three placeholders) 10. The address of the sub-processor list page and the change-notification mechanism (§8.1) 11. Whether the Japanese distributor クリューグル株式会社 (Krugle Inc.) handles user personal information; if it does, its role and the data categories must be stated and an agreement signed (§8.1) 12. Whether third-party content-safety review is enabled during the beta (§8.4)
International transfers 13. Production deployment regions (§9) 14. Provider regions; separate treatment if any are in mainland China (§9) 15. SCC modules and execution date; UK IDTA / Addendum (§9)
Retention (§10, line by line) 16. Grace period after account deletion 17. How long deleted content persists in backups 18. Session and container record retention (backend default is currently 30 days) 19. Voice and TTS cache TTL (§3.6 and §10 — two placeholders) 20. Usage and billing record retention (statutory tax and accounting period) 21. Sign-in / sign-out audit log retention 22. Server log retention (currently 30 days by default) 23. Invite redemption and device identifier retention, and whether expiry drops plaintext while retaining hashes — this is the period referred to by “a defined retention period” in 6.4(d) 24. Waitlist record clearance rule
Implementation status — verify before launch; delete anything unimplemented 25. Invite redemption forensic fields: columns exist, but as of 2026-08-25 neither client reporting nor server write logic is implemented, and nothing computes the IP hash (§3.2) 26. IP country/region code: lookup not implemented as of 2026-08-25 (§3.3) 27. Device identifier collection: not implemented on the client, no write logic on the server (§6.1). Note that the storage form (clear text) is settled — what remains open is only whether the collection ships at all
Client behaviour remediation 28. The selection toolbar exclusion list ships empty — recommend pre-populating with password managers, terminals, and banking apps, and showing a clear first-run explanation (§3.9) 29. The companion browser’s manifest carries unused camera and Bluetooth declarations — recommend removing them; if retained, they must be disclosed here (§3.9) 30. The file sync component’s global discovery and relay defaults are not explicitly disabled — recommend disabling them explicitly; if retained, disclose here (§3.9)
Cookies and user rights 31. Whether the product website and browser client adopt analytics cookies; if so, a consent banner is required (§12. The corporate website’s Google Analytics is governed by its own policy and is out of scope here) 32. When self-service account deletion and data export will ship (§11) 33. Committed response time for data subject requests (§11, §18)
Fact-checking the security claims 34. Whether a DPA is in place with every processor (or their standard DPA accepted); anything outstanding must not be stated as an unqualified commitment (§8.1) 35. The edge WAF and DDoS protection actually enabled in production (§14)
Operational tasks before launch (not marked as placeholders in the body, but required before going live)
- Set up dedicated email aliases. The only address published today is
sales@archaea-ai.com, which is what the body of this document uses. Before launch, set upprivacy@archaea-ai.com(privacy and data subject requests),legal@archaea-ai.com(legal and terms), andabuse@archaea-ai.com(abuse reports), route them to a ticketing system, and replace the address in §2, §11, and §18 here and in the Disclaimer and Terms of Use. Until those aliases actually work, they must not appear in the body text. - Verify the training opt-out with each upstream provider. The “not used for training” commitment in §5.3 depends on it: before launch, confirm provider by provider that training-related options are switched off, or that the contract prohibits using our data for model training, and keep the evidence.
- Terms links on the landing page and in the client. The Privacy / Terms links in the
manufact-wwwfooter must point to text carrying the same version number as this directory. Keep them distinct from the corporate site’sarchaea-ai.com/privacy, which covers the website only — neither should stand in for the other.
© 2026 Archaea AI, Inc. All rights reserved.