Ivan - stock.adobe.com
Singapore spells out how personal data can be used in GenAI
New advisory guidelines clarifying when organisations can scrape and reuse personal data to build GenAI models are among a raft of announcements at the inaugural Singapore Data Festival
Singapore has moved to clear up one of the thorniest questions in enterprise artificial intelligence (AI) – when and how organisations can use personal data to build and improve generative AI (GenAI) models – with new advisory guidelines that also set out the conditions under which publicly available personal data can be scraped from the web without consent.
The guidelines from the Personal Data Protection Commission (PDPC) headlined the inaugural Singapore Data Festival, which opened on 20 July 2026 with a voluntary transparency framework for GenAI chatbots, guidance on digital twins and federated learning, and a cross-border data agreement with Japan.
Opening the festival, which has grown out of the annual Personal Data Protection Week, Singapore’s minister for digital development and information Josephine Teo said organisations were asking “bigger questions” about data, particularly how it can support their AI efforts.
“Without good data, even the best systems will struggle to produce useful outcomes – garbage in, garbage out. That is why data governance matters more, not less, in the age of AI,” she said. “At its heart, data governance is about trust. Customers share data only if they trust the organisation to safeguard it properly and use it responsibly.”
Under the PDPC’s advisory guidelines on the use of personal data in GenAI, organisations developing models may rely on the “publicly available exception” in the Personal Data Protection Act (PDPA) to scrape publicly accessible personal data without consent – though they will need to assess whether data sitting behind digital barriers such as paywalls and registration requirements still qualifies as publicly available.
Where personal data collected for other purposes is repurposed to develop or improve a GenAI model, and no exception to consent applies, organisations must issue AI-specific notifications that spell out what data will be used, for what purpose, and how individuals can decline or withdraw consent.
Teo cited the example of a customer service team that wants to fine-tune a model using call recordings containing names, addresses and billing details. “We are making it clear that where personal data is used to develop or improve a GenAI model, organisations should say so plainly, rather than rely on broad descriptions that users may not notice or understand,” she said.
The guidelines, which incorporate feedback from a public consultation that drew responses from 40 organisations including Google, Meta, DBS, Singapore Airlines, HSBC and WeChat, also apportion data protection responsibilities across the AI supply chain.
Specifically, model providers must comply with PDPA obligations, paying particular attention to data retention; system providers must periodically review system-level security arrangements; and system deployers bear primary responsibility for compliance, including safeguarding data flowing through their systems – especially, the PDPC noted, for agentic AI systems.
Complementing the data rules, the Infocomm Media Development Authority (IMDA) also released transparency guidelines for GenAI chatbots – which it said are among the first of their kind in the world – calling on deployers to publish chatbot info cards. Modelled on labels found on medicine and food packaging, the cards are meant to set out in plain language what a chatbot is for, what it is not for, how user data is handled and how to report problems.
Agentic AI raises the stakes
The announcements come amid the growing adoption of agentic AI systems, which industry leaders and regulators said would test existing AI governance practices.
Speaking on a panel moderated by J. Trevor Hughes, president and CEO of the International Association of Privacy Professionals, Lee Wan Sie, IMDA’s cluster director for AI governance and safety, pointed to the authority’s model AI governance framework for agentic AI, which was released in January this year and significantly expanded in May with practical examples from early deployments.
She noted that recent tests conducted with partners in South Korea found that “if you give agents a particular objective, they’ll do all they can to achieve that objective – and in that process, accidentally release sensitive data”.
Language is another weak spot. Teo revealed that an AI safety red teaming challenge held in January, involving more than 80 experts from ASEAN member states as well as China, India, Japan and South Korea, found that “casual local phrasing could slip through the safeguards that a formal-sounding request could not”.
“This vulnerability and many others that were identified were not only technical; they were also linguistic and cultural,” she said.
Markham Cho Erickson, Google’s vice-president of government affairs and public policy, estimated that agentic AI could add between $3tn and $5tn to the global economy, but acknowledged that new use cases introduce tension into privacy, safety and security rules. “One shorthand way I like to think about it is, if something is improper or unlawful to do without agentic AI, it should be improper or unlawful to do with agentic AI,” he said.
The challenge for industry, he added, is deciding when to keep agentic interactions seamless and when to intentionally introduce friction into the process where it’s important for user safety and security at critical moments. “That’s the secret sauce,” he said.
For Nimish Panchmatia, chief data and transformation officer at DBS Bank, the technology itself is the least of his worries. AI is the easy part, he said – the hard work is in getting data, processes and controls right, and too many organisations chase what the technology can do without working out what it takes to operationalise it.
DBS is applying AI in its own governance tooling, which Panchmatia regards as the only way oversight can keep pace, particularly as chatbot interactions shift from text to voice and interventions against toxic or hallucinated outputs must happen in real time. He also called for organisations to rethink siloed oversight functions, noting that cyber security, privacy and data management have become too intertwined to be governed separately.
Simon Chesterman, vice-provost at the National University of Singapore (NUS) and AI governance and policy lead at NUS AI Institute, argued that the bigger disruption to AI governance will come from the use of live data streams – as opposed to static training datasets – by AI systems.
“We don’t need governance of datasets anymore – we need governance of data flows,” he said, warning that consent-based models are straining as AI systems begin inferring information that goes far beyond the data points users agreed to share. Users, he suggested, may need “something more like a prenup” for their relationships with increasingly human-like AI systems.
Guidance on digital twins and federated learning
Separately, the IMDA launched the Digital Twin for Enterprises Playbook, touted as Singapore’s first practical field guide to help non-ICT enterprises adopt the technology.
Among its case studies is facilities management firm Exceltec, whose digital twin consolidates data from internet of things sensors, building systems and maintenance logs across its managed sites into a single platform – eliminating manual pump inspections to save about 45 minutes of technician time daily, and opening up new revenue as the company looks to commercialise the solution to partners.
The PDPC also released a guide on federated learning, which enables collaboration without sharing original data, and updated its guide on synthetic data generation, alongside a refreshed data protection guide for ICT systems that now covers AI implementation practices.
Rounding out the announcements, the PDPC signed a memorandum of cooperation with Japan’s Personal Information Protection Commission to promote cross-border privacy rules and develop model contractual clauses, in a bid to lower compliance costs for businesses transferring data between the two countries.
With Singapore set to assume the ASEAN chairmanship next year, Teo said the city-state would work with regional partners to bring their governance approaches closer together and reduce friction for businesses. “None of us can build a trusted data ecosystem by looking only within our borders,” she said.
Read more about data governance
- AI models and agents can automate data governance policy enforcement for proactive governance in dynamic data environments. But don't forget the human element.
- AI agents will transform complex data management, optimise cloud costs and overcome the limitations of standalone generative AI, according to Gartner.
- Implementing a data governance program isn't enough. Data leaders also need to track and analyse various metrics to evaluate its effectiveness and address shortcomings.
- A strong data governance strategy enables more effective data use and helps prevent financial, legal and reputational problems. Follow these steps to develop one.
