Key takeaways

  • ✓Voice cloning tools can now replicate a person's voice from a few minutes of audio, enough to produce a convincing call from a "CFO" instructing a finance team to transfer funds.

  • ✓Business email compromise (BEC) fraud is evolving: attackers are combining cloned audio with fake emails and fabricated documents to create multi-channel attacks that are much harder to dismiss as obvious scams.

  • ✓Australian organisations are actively targeted, partly because of high average transaction values and a business culture that can make questioning a senior executive feel uncomfortable.

  • ✓The roles most exposed are not just finance staff. Executive assistants, HR managers, and anyone with access to payroll, payments, or sensitive credentials are all in scope.

  • ✓Technical controls help, but the most reliable defence is a verified call-back protocol that does not rely on the original communication channel, combined with staff who have been trained to use it without embarrassment.

How does voice cloning fraud work in practice?

A voice cloning attack follows a short, repeatable sequence. An attacker harvests audio, builds a synthetic voice model, and then calls someone inside your organisation pretending to be an executive. The whole chain can be completed in hours.

Step one: collecting the audio

Attackers do not need a recording studio. They need around 30 to 60 seconds of clean audio from a real person, and that audio is often publicly available. LinkedIn video posts, conference recordings, earnings calls, media interviews, and even internal town hall recordings that have been shared externally all serve as source material. For senior executives who have any public profile, this is rarely a barrier.

Step two: generating the voice model

Several commercially available tools can produce a convincing synthetic voice from a short audio sample. The quality of these models has improved sharply over the last two years. The clone does not just reproduce pitch and tone; it picks up rhythm, filler words, and the characteristic way a person phrases things. To an untrained ear on a phone call, with normal background noise and a degree of call anxiety, it is genuinely difficult to tell it apart from the real person.

Step three: building the pretext

The voice alone is rarely enough. Attackers research their target before making the call. They review the target's LinkedIn connections, look at what the executive has said publicly about the business, and time the call to coincide with a plausible scenario: a board meeting, a deal in progress, a regulatory deadline. The pretext typically invokes urgency and confidentiality, both of which suppress the instinct to verify.

Step four: making the call

The call itself is brief and directed. The synthetic voice claims to be the CFO, the CEO, or a senior partner, and asks a finance or operations contact to process a payment, share credentials, or approve a transfer. The target is told not to use normal channels because of the sensitivity of the matter. That instruction is the attack's most important element. It removes the safeguard that would otherwise stop it.

The instruction to bypass normal channels is the attack

Every other element, the cloned voice, the urgent pretext, the plausible timing, exists to make that one instruction feel reasonable. If your people know to treat any request to skip verification as an automatic red flag, the attack fails at the last step regardless of how convincing the voice sounds.

Some attacks pair the call with a spoofed email sent seconds before or after, so the target receives two apparently independent signals confirming the request. Two signals from two channels feels like corroboration. It is the same attacker using two vectors simultaneously.

This is the core mechanic of what is now being called a next-generation business email compromise (BEC) attack. Traditional BEC relied on a written email impersonating an executive. Voice cloning fraud adds an audio layer that most organisations have no procedure to verify.

Why are Australian organisations a target?

Australian businesses lose more to business email compromise than almost any comparable economy on a per-capita basis. The Australian Federal Police and the ACCC's Scamwatch have both flagged BEC as one of the highest-value fraud categories targeting local organisations, with losses running into the hundreds of millions of dollars annually. Voice cloning fraud is the next evolution of that same playbook.

Several factors make Australian organisations particularly attractive.

Geography and the remote-work baseline. Distributed teams are now the norm across most Australian industries, and the pandemic permanently normalised approving financial requests over video calls, voice messages, and messaging apps rather than in person. When a "CFO" rings to authorise a same-day transfer, there is no face-to-face moment to trigger doubt.

Wire transfer culture. Australia's financial system moves large sums quickly, and the business expectation of same-day settlement is routine. Fraudsters rely on urgency and speed, and our payment infrastructure accommodates both. Once funds leave via OSKO or an international wire, recovery is rare.

Time-zone isolation. When an attack targets an Australian subsidiary of a multinational, head office is asleep. The local finance team cannot easily reach anyone to verify an unusual instruction, and attackers know this. That window is often all they need.

Publicly available executive audio. LinkedIn video posts, company webinars, podcasts, and media appearances give attackers a steady supply of voice samples. A credible clone can be assembled from as little as a few minutes of clean audio, and Australian executives are increasingly visible online.

The barrier to entry has collapsed

A convincing voice clone no longer requires specialist equipment or expertise. Cloud-based tools have made it fast, cheap, and accessible to anyone willing to misuse them. The technical sophistication that once separated amateur scammers from serious threat actors is largely gone.

For finance teams and security leaders, the risk is not hypothetical. If your organisation runs wire transfers above a certain threshold, maintains overseas banking relationships, or has executives with a meaningful public profile, you are a plausible target. The question is whether your verification processes were designed with this threat in mind, or whether they were built for a world where a familiar voice on the phone was still reliable evidence of identity.

What makes voice cloning fraud harder to detect?

The short answer is that the attack is engineered to sidestep every instinct a careful person normally relies on.

Traditional business email compromise (BEC) had a tell: something in the writing felt off. The phrasing was slightly wrong, the sender address didn't quite match, or the urgency was out of character. A reasonably alert employee could spot it. Voice cloning fraud removes those signals almost entirely.

The voice itself is no longer a red flag

Modern voice synthesis tools can produce a convincing replica from as little as a few seconds of audio. Source material is rarely hard to find. Executives post video updates, appear in conference recordings, give interviews. That audio is public, indexable, and sufficient.

The replicated voice carries the right cadence, accent, and speech patterns. It hesitates in the right places. It sounds tired on a Monday morning call. For most listeners, hearing a familiar voice is the end of the verification process, not the beginning.

Caller ID gives the attack a second layer of credibility

Spoofed caller ID means the call appears to come from a number the recipient recognises: the CFO's mobile, the CEO's direct line, a number saved in the company directory. The employee sees the name on screen before they even pick up.

Combined with a convincing voice, this creates a situation where the recipient has two independent signals telling them the call is legitimate. Both are false, but the human brain treats them as corroborating evidence rather than a single point of failure.

Two signals, one fabricated source

Spoofed caller ID and a cloned voice feel like independent confirmation. They are not. Both originate from the same attacker. Treat them as a single data point, not two.

The call is designed to compress your thinking time

Attackers do not give recipients space to reflect. The scenario is almost always urgent. A payment must clear before a deadline. A deal will collapse if the transfer doesn't go through in the next hour. Legal or regulatory pressure is invoked. The emotional register is stress or controlled panic, which is harder to fake in text but perfectly replicable in voice.

That time pressure is deliberate. Scepticism requires cognitive bandwidth. When someone is under stress and believes a senior person needs something immediately, the mental space needed to question the request simply isn't there.

Real-time synthesis closes the last gap

Earlier voice synthesis tools were mainly useful for pre-recorded audio. An attacker could generate a convincing voicemail, but a live two-way conversation was technically difficult to fake convincingly.

That gap has closed. Real-time voice conversion tools now exist that can transform a live voice into a target's voice with minimal latency. This means an attacker can hold a genuine back-and-forth conversation, answer follow-up questions, respond to pushback, and maintain the persona across the full duration of the call.

An employee who senses something is wrong and asks a clarifying question no longer gets the hesitation or deflection that might have revealed the fraud. They get a plausible answer, in the right voice, almost immediately.

Social engineering does the rest

The technical elements are only half the attack. The other half is knowing which request to make, to which person, at which moment in the business cycle. Attackers research their targets. They know when a company is closing a deal, processing payroll, or managing a supplier relationship under pressure.

A well-researched approach means the request itself doesn't feel strange. "We need to move the supplier payment early this quarter" lands differently on someone who knows a supplier renegotiation is in progress than it does in isolation. The attack is contextually plausible because the attacker has done homework.

This is what makes voice cloning fraud qualitatively different from earlier forms of BEC. The technical barrier to a convincing impersonation has dropped dramatically. The social and contextual research required is exactly the kind of work that motivated, well-resourced criminal groups have always been willing to do.

Which roles are most exposed?

Voice cloning fraud is targeted, not random. Attackers research an organisation before they call, and they call the person most likely to act without questioning. Four roles come up repeatedly.

Finance approvers. Accounts payable staff and finance managers are the primary target. They handle payment instructions daily, they are trained to process requests efficiently, and they often receive calls that mimic a CFO or managing director asking for an urgent wire transfer. The combination of authority impersonation and time pressure is designed to bypass normal process. See our CFO's guide to AI deepfake fraud for a deeper look at how these attacks are structured.

Executive assistants and personal assistants. EAs and PAs hold significant operational access. They book travel, manage credentials, approve invoices on behalf of executives, and often have authority to communicate decisions to third parties. An attacker who can convincingly impersonate a CEO has a ready-made path through the EA to almost any part of the business.

IT administrators. System access and credential resets are valuable targets in their own right. An IT admin who receives a call from someone who sounds like the CTO asking for an urgent account unlock or a VPN credential bypass is in a difficult position. The request sounds legitimate, the voice matches, and the situation feels like it demands immediate action.

Board members and senior executives. Executives are impersonated most often, but they are also targeted directly. A board member receiving what sounds like a trusted peer requesting a confidential document, or a CEO getting a call that appears to come from their CFO before a major transaction, represents a different attack vector. At this level, the social engineering is usually more sophisticated and the potential loss is higher.

The pattern across all four

Every high-exposure role combines some form of authority (to approve, to access, or to instruct) with a reasonable expectation of urgent, out-of-process requests. That combination is exactly what attackers are engineering for.

One thing worth noting is that exposure is not purely about seniority. A mid-level accounts payable officer with authority to approve transfers under a certain threshold can be just as valuable a target as a CFO, especially if their verification habits are weaker.

How do you build a defence that actually holds?

Technology alone won't solve this. Voice cloning fraud works because it exploits process gaps and human instincts, not just security tooling. The most effective defences combine a few hard rules with regular practice so those rules hold up under pressure.

Establish a callback policy for every payment instruction

The single most effective control is a mandatory out-of-band verification step for any instruction to transfer funds, change payment details, or create a new payee. "Out-of-band" means using a different communication channel from the one that delivered the request. If the instruction came by phone, verify it by email or a message sent through your internal system. If it came by email, call back on a number from your internal directory, not one supplied in the message.

This sounds obvious. Teams skip it because the caller sounds urgent, authoritative, or both. The policy needs to be unconditional: no dollar threshold below which it does not apply, no seniority level that bypasses it.

The rule that stops most attacks

A callback to a verified number, on a different channel from the original request, breaks the attack chain at its weakest point. Make it non-negotiable, regardless of who is asking or how much urgency they convey.

Agree on a shared verification word or phrase

Some organisations use a simple pre-agreed code word between executives and their finance or operations contacts. When a high-stakes instruction arrives and something feels off, the receiving person asks for the word. A real executive gives it immediately. An attacker using cloned audio cannot, because the word was never in any recording.

This works best for small teams where the same people regularly authorise large transactions. It requires almost no technology and costs nothing to implement.

Treat urgency as a red flag

Legitimate payment instructions rarely require someone to act in the next ten minutes. If a caller is pushing hard for immediate action, that pressure itself should trigger the verification step, not bypass it. Train your finance and operations staff to recognise urgency as a manipulation technique and to slow down in response to it, not speed up.

A useful framing for training: any real CFO or CEO would rather wait an extra hour for a payment than have funds transferred fraudulently. If the person on the phone won't accept a callback, that is your answer.

Run scenario-based awareness training

Policies on paper do not hold up if staff have never practised them under simulated pressure. Scenario-based training, where employees work through realistic voice phishing (vishing) situations, builds the muscle memory to follow the protocol when adrenaline is running. This is different from a compliance module with a quiz at the end.

Effective training for this threat covers:

  • What a voice cloning attack sounds like, including audio examples

  • The specific manipulation techniques attackers use (urgency, authority, partial information to seem credible)

  • The exact steps the employee should take, practised until they are automatic

  • What to do after a suspicious call, including how to report it internally

Better People's AI scam and deepfake awareness workshop covers exactly this ground, with content tailored to the roles most likely to be targeted. For a broader look at governance and policy framing, the practical AI governance framework for mid-sized enterprises is a useful companion.

Review your authorisation structure

Many organisations have single-point approval for significant transactions. One person gets the call, one person approves the transfer. Dual authorisation, where two people must independently confirm a transaction above a set threshold, dramatically reduces the chance that a single successful impersonation leads to a loss. It also means an attacker would need to clone two voices and reach two separate people, both of whom follow a compromised process.

This is a structural change, not just a training one. Finance teams sometimes resist it for speed reasons. The conversation about that trade-off is worth having explicitly before an incident makes it unavoidable.

Frequently asked questions

How can you tell if a voice call is a cloned recording or a live person?

There is no reliable way to detect a cloned voice in real time with your ear alone. Modern voice synthesis tools can replicate tone, pace, and even background noise well enough to fool a trained listener. The practical defence is not detection but verification: any call requesting a financial action should trigger a callback to a known, pre-registered number, regardless of how convincing the caller sounds. Relying on audio cues such as slight robotic quality or unnatural pauses is becoming less useful as the technology improves.

What is the difference between a deepfake and a voice clone?

A voice clone replicates only the audio of a person, typically from a few minutes of recorded speech. A deepfake refers to a synthetic video or image, often combining a fabricated face with cloned audio. In financial fraud, voice clones are currently more common because they are faster and cheaper to produce and work over a standard phone call. Deepfakes require more processing power and are more likely to appear in video calls. Both belong to the same family of AI-generated impersonation risk, and your verification protocols should cover both. The CFO's guide to AI deepfake fraud covers the video-based variant in more detail.

Who is legally liable if a payment is made following a voice cloning attack?

Australian law does not offer a clean answer here, and liability often depends on whether the organisation followed its own documented controls. If a finance team member bypassed an existing dual-authorisation policy because a caller "sounded like the CFO", the organisation may struggle to recover the funds or claim on insurance. Cyber insurance policies vary significantly in how they treat social engineering losses, and some require specific controls to be in place as a condition of cover. Legal advice specific to your policy and circumstances is worth seeking before an incident occurs, not after.

How often should staff receive training on voice cloning and impersonation fraud?

Annual training is a floor, not a standard. The attack methods are evolving quickly enough that a single annual session leaves long gaps in awareness. A more effective approach is short, targeted refreshers every quarter tied to new attack patterns, combined with simulated phone-based social engineering exercises. Roles with direct payment authority or access to executive schedules warrant more frequent touchpoints than the general workforce. The article on who needs AI scam awareness training, and how often explores frequency and targeting in more depth.

Can a verification code or passphrase actually stop these attacks?

Yes, and it is one of the most practical controls available. A pre-shared passphrase between executives and finance staff works because an attacker cannot easily know it, regardless of how well they replicate a voice. The passphrase needs to be changed periodically, stored securely, and never communicated over email or messaging platforms that may themselves be compromised. The limitation is consistency: the control only works if every person in the chain uses it every time, without exceptions made for urgency or seniority. An attacker who knows your culture well enough may engineer a scenario specifically designed to make the pause for verification feel awkward.

Is your team ready for the call that sounds real?

Voice cloning fraud is not a future risk. Australian organisations have already lost money to it, and the attacks are getting cheaper and easier to run. The question is not whether your finance or executive team will encounter one, but whether they will recognise it in the moment.

That recognition does not come from a policy document sitting in a SharePoint folder. It comes from practice: hearing what a convincing impersonation sounds like, running through a verification step under mild pressure, and knowing exactly when to stop a transaction and make a callback.

Better People's AI Scam and Deepfake Awareness workshop is built for finance teams, executive assistants, and the leaders who approve payments. It covers how voice cloning and deepfake attacks are constructed, what the live warning signs look like, and how to apply a verification protocol without stalling legitimate work.

Want to run a session for your finance or executive team?

We will walk you through what the session covers, how long it runs, and whether it fits your team's context. No commitment required.

See the AI scam awareness workshop →

For a broader look at how AI is reshaping fraud risk for financial leaders, the CFO's guide to AI deepfake fraud covers the wider threat picture alongside the governance decisions that sit behind it.