Obiguard releases obi-lookout, a free open model that finds sensitive data in English and Malay text
A small open model, free under Apache 2.0, that finds personal and business-sensitive data in English, Malay and mixed text and runs on an ordinary CPU.
KUALA LUMPUR — Obiguard, the AI risk governance company, today announced the release of obi-lookout, a small open model that finds personal and business-sensitive data in text, so it can be masked or blocked before it reaches a language model or a log. It is free to use under the Apache 2.0 licence and is published on Hugging Face, with a quick-start guide on GitHub.
obi-lookout reads text and returns where each sensitive value is and what kind it is: names, ID numbers, phone numbers, email addresses, addresses, dates of birth, bank account numbers, money amounts, secrets such as API keys, internal hostnames, and client or project names. It was built for English, Malay and text that mixes the two, and handles Malaysian and Singapore formats. It is a detector, not a chat model: it returns positions and labels, not generated text.
The model is small enough to run on an ordinary CPU, with no GPU, and it works offline once downloaded, so the text it checks never has to leave an organisation’s own network. The quick-start repository shows how to use it from Python, from the command line, and as a Docker server that speaks the OpenAI API format, so tools and gateways that already call an OpenAI-style endpoint can call it too.
On a test set of 1,569 documents written from generated data, obi-lookout scored an overall F1 of 0.927, counting a hit only when the start, end and kind of a value were all exact. On documents containing nothing sensitive, 7.1% had at least one false alarm. Obiguard says plainly that the test set is synthetic and real documents will differ, that the model is not a guarantee that every sensitive value is found, and that it is weaker on organisation and client names and on text that is not prose, such as shell sessions and logs. The evaluation data is published so others can reproduce the numbers.
obi-lookout is built on the open-source GLiNER2 framework and fine-tuned from Fastino’s GLiNER2 privacy model. It was trained only on generated data, with no customer data and no real personal data. The model card lists its licences and attributions.
Security and engineering teams can use it to mask sensitive values before text is sent to an AI service or stored in logs, and to check what prompts and records contain. Because no detector is complete, it is meant to be one control among several. Obiguard also builds an AI governance platform that enforces policy on AI model calls and keeps an audit trail.
obi-lookout v0.1.0 is available now on Hugging Face. The quick-start guide takes a few minutes to run.
Obiguard is an AI risk governance platform founded in Malaysia in 2024 and backed by Antler. It helps organisations inspect AI usage, enforce policy, and keep the audit evidence regulators expect. Learn more at obiguard.ai/about.
Media contact: [email protected]