OpenAI "Project Lily" Exposed: Contractors Can Access Users' Chat Histories
nashnova research
OpenAI's internal initiative 'Project Lily' has been exposed — hundreds of contractors are reading real users' ChatGPT conversations and scoring AI replies, putting the privacy of over 900 million users under scrutiny.
What exactly are the contractors reading?
Project Lily follows a three-step workflow: read a real user's prompt, summarize the user's intent, then score four system-generated replies on a 1-to-7 scale with written commentary.
A core training goal is to reduce ChatGPT's "sycophantic" personality — OpenAI already faces multiple lawsuits alleging GPT-4o's excessive agreeableness contributed to self-harm incidents.
This means → your full ChatGPT conversation, context included, may be reviewed line by line by an outsourced team you have never heard of.
Does the privacy filter actually work?
OpenAI says prompts pass through an internal "privacy filter" model before reaching reviewers, stripping personal information. Contractors cannot see usernames.
But the company itself concedes the filter "may miss rare identifiers or ambiguous private expressions" and can "over-redact or under-redact" when context is limited.
Crucially, prompts sometimes carry a "user memory summary" above them — an outline of the user's past chatbot usage, including details like location. In plain terms = even with the name removed, the conversation itself can still sketch out who you are.
Were users told?
When asked whether users were explicitly informed their prompts would undergo human review, OpenAI gave no clear answer.
Its website previously stated human review might occur for policy violations or safety risks — model-optimization review was not part of the public disclosure.
Only after the story broke did OpenAI point to existing language about human review for "improving model performance" and update its help page with detailed opt-out instructions. This means → the disclosure was completed under pressure, not proactively.
How do you opt out?
Users can toggle off "Improve the model for everyone" in settings; once off, chat logs are no longer used for training.
But the toggle is on by default for Free, Plus, and Pro accounts. Users must turn it off manually, and the change applies only to new conversations — past logs cannot be retroactively excluded.
Enterprise, Business, and Education accounts have the toggle off by default. In plain terms = the vast majority of everyday users have been opted in since the day they signed up, unless they switch it off themselves.
Is the outsourcing chain itself secure?
Review positions are staffed through recruitment agency Crossing Hurdles and paid by AI training firm Mercor at over $50 per hour.
Mercor suffered a major data breach in April 2026; Meta has since terminated its relationship with the company.
This means → the outsourcing chain handling user data has its own unclean security record.
Is OpenAI the only one doing this?
No. Google Gemini's disclaimer explicitly mentions human reviewers may examine some saved chats; Anthropic confirms it uses human review of select conversations to improve Claude.
Michal Luria, senior researcher at the Center for Democracy & Technology, argues human review may be essential for chatbot safety — but the mechanism clashes with users' sense of privacy in a chat interface.
This reflects a structural, industry-wide tension: models need human feedback to become safer, yet the process of collecting that feedback erodes user trust. Whether more transparent disclosure can resolve this contradiction will be a central focus for regulators.
市场有风险,内容仅供研究参考,不构成投资建议。
