Now is the time to think about AI privacy

Now is the time to think about AI privacy

Your chatbot logs are always at risk of exposure. Here are some practical tips to increase your AI privacy.

This week, thanks to blockbuster reporting on 404 media, we have some insight into the dire state of privacy when using AI chatbots:

OpenAI is hiring hundreds of contractors who read a massive stream of real users’ ChatGPT prompts, with the prompts sometimes including sensitive personal information, 404 Media has learned. The prompts these people review can include whole conversations between users and the chatbot, conversations that most of ChatGPT’s more than 900 million users probably don’t realize may be read by actual people.

Behind every prompt, picture, or snippet of code sent off to ChatGPT, there is at least some possibility that information will not only be stored and archived for later retrieval (think cops and lawyers), but also that contractors will review these AI conversations in hopes of improving the models.

This isn’t a surprise to most tech folks, but consumers deserve to know that this is the status quo for any and all AI apps.

Your intimate therapy sessions, the photos of that suspicious mole, and the wording of your letter of resignation to your boss – if you’ve put any of that into a closed-source model or app, it exists on someone else’s server.

OpenAI’s product is the focus of the reporting, but the same is true if you’re using Claude, Gemini, Grok, or Muse by Meta.

The Coming Storm

For policy-conscious folks, the hoards of intimate data and confidential information contained in these apps will continue to be a public policy issue that will shape the future of AI privacy.

There will be demands by courts, police, intelligence, and maybe even employers to have some visibility or access into our chat logs. Chatbot logs are already a common trope brought up in courtroom and police dramas on television, even more than Google search history (cue to The Rookie!).

Also, there are already a handful of court cases that relate to whether this information should be protected under the Fourth Amendment, and what kind of warrants or court orders agencies and lawyers will need before accessing this information.

In NYT v. OpenAI, regarding the use of newspaper articles for training LLMs, the makers of ChatGPT were already compelled by the judge to produce nearly 20 million chat logs of individual users. Those logs are supposedly “de-identified,” but we know they’ll contain sensitive information.

This is where we must began to ask smart questions and have ready-made answers that will protect consumer privacy.

What obligations and liabilities will be placed on AI companies for sharing the context and information of chats? What about when it relates to crime, terrorism, corporate espionage, or any other type of suspected wrongdoing? The implications of these questions are far-reaching for individual privacy, and are no doubt already a concern for many people working at AI firms across the country.

The horrific shooting in Tumbler Ridge, British Columbia earlier this year, in which eight people were killed by a lone gunman who used ChatGPT to plot some elements of the killings, has already been seized upon by Canadian politicians as evidence that “AI Tech” needs more checks and balances. And that’s despite the many failures by both police and the medical system that have nothing to do with technology.

The pressure to monitor, surveil, and moderate your AI chatbot conversations will only increase. The same movement that sought some kind of “lawful access” into encrypted messaging, phone calls, and internet traffic will soon set their sights on what you asked Claude or ChatGPT.

An Open Way Out

Beyond the societal conversation that will inevitably erupt around the privacy of our AI use, the best way to personally protect yourself relies on using the technology itself.

This is the key selling advantage of using open-weight and open-source models that are also available for consumers to use.

As I wrote this week in Daily Wire, open models allow for much more flexibility and privacy to guard against your data leaking into court cases or into the hands of snoops from every corner:

Closed-source models are black-boxed technologies, centrally hosted and controlled by AI companies. Open models, on the other hand, allow users to download them, tinker with context and reasoning, and privately host them on their own machines. Once a developer ships an open model, there are no reasonable measures in place to enforce compliance with users or rein in its use.

The recent purchase of the open model library Hugging Face by top chipmaker NVIDIA was a huge vote of confidence in open-weight models and has given a boost to thousands of open-source AI projects that use the platform to distribute their technology. This kind of deal would let open-source models advance and compete with frontier tech. But only if we allow it.

To that end, we want to offer some practical tips to help users improve their AI privacy:

Start segmenting your AI use

You should think of your AI use as separate buckets or folders. If you have an AI subscription for work purposes, such as a Claude or GPT, then use them only for that purpose.

At least for enterprise and corporate subscriptions, both Anthropic and OpenAI maintain “Zero Data Retention” policies that supposedly keep prompts, data, and results from being used to train their models. We have no way to prove this, either in code or direct inspection, so be cautious. But at least based on their terms of service, enterprise accounts are supposedly more private than individual accounts, paid or not.

If you use chatbots primarily for financial conversations or budgeting, seek out solutions that focus on that specialty. The same applies to healthcare.

Be aware of Identifying Information

Before uploading your entire bank statements or pictures of your loved ones, practice discretion. AI agents don’t need your full social security number or newborn pictures in order to provide solid advice.

You can crop, selectively edit, or just use screenshots that provide enough information to answer your query. Anything you put into a chatbot will be accessible at some point, and will always be vulnerable to leaks and hacks. Stay vigilant.

Start Using Encrypted AI services

There are a growing number of AI companies offering private and encrypted inference that can demonstrably not be accessed by anyone other than you.

Apps and API services like Maple.ai and Tinfoil are run in Trusted Execution Environments (aka “TEEs”), which are encrypted servers. The same for Brave’s AI apps. Additional products like Venice.ai have their own trade-offs when it comes to privacy and encryption and warrant a look. Start exploring them.

You also have great options from ProtonMail in Luma, or even DuckDuckGo’s Duck.ai sevice.

For the most intimate and private conversations, these can give a user much better piece of mind.

Go local

The only verifiable private way to use AI inference, agents, and chatbots is to self-host models on computers you control. Model libraries like Hugging Face are making that easier to do, as well as the routing companies that are helping users set up their own local models on shared servers.

Open models from companies like Meta, NVIDIA, and even Google are giving users an opportunity to run models at home, and the technology is only getting better as competition improves. AI assistant ecosystems like Hermes Agent and OpenClaw give users the ability to change their API keys or models to be more private if they wish, and they’re only getting better.

The true cost of running frontier-like models at home is still very high, and that should be acknowledged. But options like LM Studio are allowing people to run models on their phones, as well as dozens of other companies coming up.

Find a way

Regardless of the option you want to take, you should start thinking about AI privacy. It’s not too late.

Published on The Upgrade (archive #1, #2)