OpenAI Unveils Privacy Filter for 128K Context with Eight Recognition Modes
OpenAI has introduced Privacy Filter, a state-of-the-art Personal Identifiable Information (PII) anonymization model. Now available under the Apache 2.0 license on Hugging Face and GitHub, this tool empowers developers with a local, highly customizable solution for robust privacy protection.
Advanced Semantic Understanding Beyond Rule-Based Matching
Unlike traditional rule-based tools, Privacy Filter leverages deep language comprehension to accurately identify sensitive data within unstructured text based on context. This approach effectively obscures private information while preserving maximum utility in public content.

Lightweight MoE Architecture for Superior Efficiency
The model’s technical design prioritizes both flexibility and performance:
Mixture of Experts (MoE) Design: With a total parameter count of 1.5 billion, only approximately 50 million parameters are activated per inference. This efficiency enables smooth operation on resource-constrained edge devices, including laptops and web browsers.
Extended Context Support: Featuring a 128,000 Token context window, the model utilizes a bidirectional token classification architecture combined with a constrained Viterbi algorithm to ensure accuracy and coherence when processing long documents.
High Accuracy Recognition: In the updated PII-Masking-300k benchmark, the model achieved an F1 score of 97.43%, boasting a recall rate of 98.08%.
Comprehensive Privacy Classification System
Privacy Filter accurately identifies and labels eight core categories of sensitive information:
Basic Identity: Names, addresses, email addresses, and phone numbers.
Online Assets: URL links.
Financial Security: Account details, including bank and credit card numbers.
Confidential Credentials: Passwords and API keys.
Time-Sensitive: Date information.
Application Scenarios: A "Local Firewall" for Cloud LLMs
OpenAI positions this tool as a pre-filter layer. By processing text locally for PII detection and anonymization before sending it to cloud-based large language models, this "data stays on device" strategy significantly reduces the risk of users inadvertently exposing private information to AI services.
Related article
NewCore raises $66M to equip AI agents with identities as they transition into employee roles
Cybersecurity startup NewCore has officially launched from stealth, securing $66 million in funding on Monday. The company aims to address a critical challenge that many organizations will soon encounter as they integrate AI agents: how to authentica
Microsoft and Mistral Ink Multi-Billion-Dollar European AI Partnership
At the core of this collaboration lies a multi-billion-dollar pact designed to bolster artificial intelligence infrastructure across Europe. Credit: MicrosoftMicrosoft and Mistral, a leading French artificial intelligence firm, have unveiled a multi-
Microsoft exec calls AI data scraping 'biggest theft of labor in history,' unredacted filings show
Unsealed documents from the three-year-old copyright lawsuit filed by The New York Times against OpenAI and Microsoft reveal an admission that AI scraping constitutes theft and poses a severe threat to media outlets.According to the lawsuit, a senior
Related Special Topic Recommendations
Comments (0)
0/500
OpenAI has introduced Privacy Filter, a state-of-the-art Personal Identifiable Information (PII) anonymization model. Now available under the Apache 2.0 license on Hugging Face and GitHub, this tool empowers developers with a local, highly customizable solution for robust privacy protection.
Advanced Semantic Understanding Beyond Rule-Based Matching
Unlike traditional rule-based tools, Privacy Filter leverages deep language comprehension to accurately identify sensitive data within unstructured text based on context. This approach effectively obscures private information while preserving maximum utility in public content.

Lightweight MoE Architecture for Superior Efficiency
The model’s technical design prioritizes both flexibility and performance:
Mixture of Experts (MoE) Design: With a total parameter count of 1.5 billion, only approximately 50 million parameters are activated per inference. This efficiency enables smooth operation on resource-constrained edge devices, including laptops and web browsers.
Extended Context Support: Featuring a 128,000 Token context window, the model utilizes a bidirectional token classification architecture combined with a constrained Viterbi algorithm to ensure accuracy and coherence when processing long documents.
High Accuracy Recognition: In the updated PII-Masking-300k benchmark, the model achieved an F1 score of 97.43%, boasting a recall rate of 98.08%.
Comprehensive Privacy Classification System
Privacy Filter accurately identifies and labels eight core categories of sensitive information:
Basic Identity: Names, addresses, email addresses, and phone numbers.
Online Assets: URL links.
Financial Security: Account details, including bank and credit card numbers.
Confidential Credentials: Passwords and API keys.
Time-Sensitive: Date information.
Application Scenarios: A "Local Firewall" for Cloud LLMs
OpenAI positions this tool as a pre-filter layer. By processing text locally for PII detection and anonymization before sending it to cloud-based large language models, this "data stays on device" strategy significantly reduces the risk of users inadvertently exposing private information to AI services.
NewCore raises $66M to equip AI agents with identities as they transition into employee roles
Cybersecurity startup NewCore has officially launched from stealth, securing $66 million in funding on Monday. The company aims to address a critical challenge that many organizations will soon encounter as they integrate AI agents: how to authentica
Microsoft and Mistral Ink Multi-Billion-Dollar European AI Partnership
At the core of this collaboration lies a multi-billion-dollar pact designed to bolster artificial intelligence infrastructure across Europe. Credit: MicrosoftMicrosoft and Mistral, a leading French artificial intelligence firm, have unveiled a multi-
Microsoft exec calls AI data scraping 'biggest theft of labor in history,' unredacted filings show
Unsealed documents from the three-year-old copyright lawsuit filed by The New York Times against OpenAI and Microsoft reveal an admission that AI scraping constitutes theft and poses a severe threat to media outlets.According to the lawsuit, a senior





Home






