WhatsApp Begins Testing AI Feature That Flags Scam Messages Before Users Reply

WhatsApp has started rolling out a beta version of “Scam Alert,” a feature that runs a machine learning model directly on a user’s phone to warn them when an incoming message from a stranger resembles a known scam. Meta confirmed the limited rollout began on August 12, 2026, and the company is currently testing it with security researchers in WhatsApp’s Bug Bounty community.

Unlike earlier WhatsApp safety tools that relied on account behavior or contact metadata, Scam Alert is the platform’s first feature to analyze the actual content and structure of a conversation for fraud signals — while keeping that analysis confined to the device itself.

What Happened

Meta shared a preview of Scam Alert as a system designed to warn WhatsApp users about potential scam messages and contacts without compromising their privacy. According to the company, the feature is opt-in: users who enable it get a machine learning model downloaded to their phone, which then checks incoming messages from non-contacts to see if they match patterns seen in known scams. The model relies on probabilistic classification based on conversational structure and linguistic signals rather than keyword matching, meaning it looks at how a conversation is built and phrased, not just specific trigger words.

If the model flags a message, only the recipient sees a warning inside the chat, and the sender is not notified. From there, the user can block the contact, report the message, ignore the warning and continue the conversation, or mark the chat as trusted. Marking a chat as trusted removes the warning and stops Scam Alert from flagging that conversation again, and users who do so can separately opt in to share the last five messages they received with WhatsApp to help improve the model’s accuracy.

How Scam Alert Works

FeatureDetails
StatusLimited beta, launched August 12, 2026
AvailabilityOptional; must be enabled by the user
Who it scansMessages from non-contacts only
Processing locationOn-device; no message content sent to Meta
Training dataScam conversations previously reported by users
Detection methodLinguistic and conversational-structure pattern matching
Visibility of warningShown only to the recipient, not the sender
User actions availableBlock, report, ignore, or mark chat as trusted
Can be disabledYes, at any time

Privacy and Security Safeguards

Because Scam Alert inspects message content, even locally, Meta has built additional infrastructure to reassure users and researchers that the feature doesn’t quietly reopen a window into encrypted chats. The company says the only information that reaches its servers is aggregate, anonymized data — such as how many warnings were shown and what action users took — and that this data is processed through what WhatsApp describes as a confidential federated analytics pipeline built on Trusted Execution Environments, specifically confidential virtual machines, with differential privacy noise applied before any data reaches Meta’s servers.

WhatsApp has also addressed a specific concern raised by security researchers: the risk that a malicious or compromised model could be pushed to a single targeted individual. To guard against this, the company says every model release must be logged on a third-party, append-only transparency ledger before it can be distributed, and that Cloudflare, not Meta, holds the digital signing keys used to authenticate those records, with each device checking the ledger before running a downloaded model. Some outlets have framed the design as a response to a broader policy debate: for several years, regulators in the EU and UK have pressed Meta, Signal, and other encrypted-messaging providers to enable content scanning, arguing that encryption and platform safety are incompatible — a premise Scam Alert’s on-device architecture is intended to challenge.

It’s worth noting that these privacy claims and the technical architecture come from Meta’s own announcement and its engineering blog; they have not yet been independently audited or verified by outside researchers, and the feature remains in a restricted beta rather than a public release.

Not WhatsApp’s First Move Against Scammers This Year

Scam Alert is the latest in a series of anti-fraud features WhatsApp has introduced through 2026, each targeting a different stage of a scam attempt.

DateFeatureWhat It Does
March 2026Device-linking warningsFlags suspicious attempts to link a new device to a user’s account, aimed at scams like fake “talent competition” voting schemes
June 2026Unfamiliar-number warning screenShows a screen before a chat opens with a number not in the user’s contacts, listing signals like country code and mutual groups
August 2026Scam Alert (beta)Uses on-device AI to analyze message content and structure for scam patterns

The June feature works differently from Scam Alert: rather than reading message content, it checks the phone number itself before a chat opens, showing a screen that asks the user to confirm whether they trust the sender if no signs of an existing trusted relationship are found. Meta has also disclosed enforcement numbers behind these efforts. In its March announcement, the company said it removed more than 159 million scam ads in 2025, 92% of which were taken down before anyone reported them. Separately, Meta has said it disabled 6.8 million scam-related WhatsApp accounts in the first half of 2025, many of them tied to operations in Southeast Asia.

What Remains Uncertain

Meta has not published a timeline for when Scam Alert will exit beta or which countries and app versions will receive it first. The company also has not disclosed detection accuracy figures, false-positive rates, or how the feature performs against scams generated with AI tools — an area some coverage has flagged as an evolving threat the feature will need to keep pace with. Because the beta is currently limited to participants in WhatsApp’s Bug Bounty program, most users will not be able to test it yet.

Pakistan Context

WhatsApp is one of the most widely used communication platforms in Pakistan, and the Pakistan Telecommunication Authority (PTA) has issued repeated advisories this year about scams that travel through the app. In one advisory, the regulator warned citizens about fraudsters impersonating reputable companies using stolen names, logos, and branding to appear legitimate, often targeting victims through fake job offers, urging the public to “Think Before You Click.” Separately, PTA has cautioned against messages impersonating its own identity, and warned users receiving spoofed calls and texts not to share OTPs, CNIC numbers, or banking details with unknown callers.

Pakistan’s telecom regulator has also pursued a parallel, unrelated policy: restricting WhatsApp accounts tied to unregistered or inactive SIM cards, as part of a broader push on digital identity verification. That effort is separate from Meta’s content-based Scam Alert feature and addresses a different vector — account takeover through SIM fraud rather than incoming scam messages.

Financial-sector players have also responded to WhatsApp-based fraud independently of Meta’s product changes. JazzCash, one of the country’s largest mobile wallet providers, said it has introduced multi-factor authentication, transaction alerts, and real-time fraud monitoring in response to the rise in WhatsApp-based financial scams. Pakistani users can report cybercrime through the FIA’s cybercrime helpline (1799) or financial fraud through the State Bank of Pakistan’s helpline. Meta has not said whether or when Scam Alert’s beta will expand to Pakistan.

What It Means for Users

For now, Scam Alert changes little for the average WhatsApp user, since it isn’t broadly available. But it signals a shift in how Meta is approaching fraud prevention on an end-to-end encrypted platform: rather than only reacting to reported scams after the fact, the company is testing content-level detection that runs entirely on the device. If the beta expands, users could gain an additional layer of warning before engaging with unfamiliar senders, on top of existing tools like the unfamiliar-number screen and device-linking alerts. Users should continue treating unsolicited messages from unknown numbers — particularly those involving job offers, investment pitches, or requests for OTPs — with caution, regardless of whether Scam Alert is active on their account.

What Happens Next

Meta has said it plans to continue testing Scam Alert with its Bug Bounty researchers before any wider release, and that both the model itself and the federated analytics pipeline are included in the scope of its bug bounty program. No date has been given for a public rollout.

Leave a Reply

Your email address will not be published. Required fields are marked *