June Offer Every MAX plan gets a fully custom-built system Free custom system worth $1,500-$10,000 · worth $1,500-$10,000

Voice cloning where the hard part is the timing, not the voice

Making a voice that sounds like you is a solved problem. Getting a note in that voice to the right person at the moment it will actually land, without you generating a file and uploading it, is not.

Clone your own voice once. The conversation decides when it speaks.

Get Started

The short, checkable version

Five statements from this page that stand on their own, with where each one comes from. If you are comparing tools or writing about this, these are the ones worth checking.

  • ManyChat cannot deliver a dynamically generated voice note into an Instagram DM. Its documentation states that sending audio through Dynamic Block responses or the API is unsupported on Instagram and WhatsApp, and available only on Messenger and Telegram.

    ManyChat, Response Reference for Instagram, WhatsApp and Telegram Automation

  • Audio published to Instagram through ManyChat renders as a standard audio player inside the message rather than as a native Instagram voice note.

    ManyChat, Media guidelines for Facebook Messenger, WhatsApp and Instagram automations

  • In Inflowave, a voice note is assigned to a conversation phase and released on the number of exchanges that have actually occurred, so an introduction note cannot fire late in a thread and a closing note cannot fire on the second message.

    Inflowave product

  • Inflowave requires that a cloned voice is the account holder own voice. Cloning another person voice without their permission is prohibited, independently of local law.

    Inflowave policy

  • Any voice note in Inflowave can be a real recording instead of generated speech, chosen per note rather than as an account-wide setting.

    Inflowave product

Competitor claims taken from each vendor own published documentation and checked on 30 July 2026. Products in this category change often, so verify anything that decides a purchase.

A voice note at the wrong moment is worse than none

Somebody sends you one message. Thirty seconds later a warm, personal voice note arrives from a person they have never spoken to. Nobody believes it. It reads as an automation, because it is one, and now every message after it is being read with suspicion.

That is the whole failure mode of voice in sales, and it is a timing problem rather than an audio problem. So notes here are not attached to a step in a script. They are attached to a phase of the conversation, and a phase is defined by how much conversation has actually happened.

Each phase has a range of exchanges it covers and its own note. The introduction note cannot fire on message forty, and the closing note cannot fire on message two, no matter what order things happened in or how the thread went sideways.

Introduction

The early exchanges. The job of a note here is only to prove there is a person on the other end, so it should be short and it should not sell anything. This is the one most worth recording for real rather than generating.

Warm

Once there has been a genuine back and forth. By now you know something specific about them, so this is where a generated note earns its place: it can say the specific thing, which is the difference between personal and personalised.

Closing

Deep into the conversation, where the remaining distance is usually hesitation rather than information. Tone does more than wording here, which is exactly what a voice note is for.

And one more, which is the best of them

The conversation can be held when a known objection appears, and a voice note written for that objection used in place of a text reply. If you only ever set up one voice note, make it this one.

The reason is that an objection is the exact moment a templated reply is most obviously templated. Everybody has read the paragraph that begins by saying they completely understand. Hearing an actual person say the same thing, unhurried, does something the paragraph cannot, and it is the only part of a sales conversation where that is reliably true.

The stack everybody tries, and why it does not work

The obvious build is a voice tool for the audio, an automation platform to move it, and a DM tool to send it. People spend a weekend on this. On Instagram it does not work, and the reason is worth knowing before you spend the weekend.

ManyChat can send audio on Instagram when the clip is a static file placed in a flow. Their own documentation is clear that sending audio through dynamic responses or the API is not supported on Instagram or WhatsApp, only on Messenger and Telegram. Generating a different note per lead and delivering it IS the dynamic path. So the interesting half, the personal half, is the half that does not arrive.

What you are left with is one static clip sent to everybody, which is a worse version of a text message, because at least a text message does not pretend to be personal.

Here the clone, the timing and the thread are the same system, so there is nothing to move between them. That is the entire reason this feature exists, and it is a smaller claim than the one the voice tools make. They are better at voice. We are the part that comes after.

It has to be your own voice

People ask whether voice cloning is legal, and most pages in this category answer by changing the subject. So, plainly: cloning your own voice is fine. Cloning the voice of another person without permission can be illegal, and the law has moved fast, with several US states now carrying specific statutes on synthetic voice and likeness.

Our rule is simpler than the law and easier to remember. It must be your own voice. Not a celebrity, not a competitor, not a colleague who has not asked for it, and not a client whose voice you think would sell better than yours.

The test that settles it in every ambiguous case: would you be comfortable telling the person whose voice it is exactly what you are doing with it? If the answer takes more than a second, do not clone it.

There is a commercial argument for the same position, which is that the thing making a voice note work is that it is genuinely you. A cloned voice belonging to somebody else is not a shortcut to trust, it is a liability with a good accent, and it only has to be discovered once.

Inflowave create-voice-clone screen with two audio samples uploaded, their file sizes and durations listed against a recommended sixty to a hundred and eighty seconds, and the audio requirements shown above
One good sample, recorded once

How it is delivered, not just what it says

The same sentence in your voice can be steady or animated, and the right choice depends on where you are in the conversation. Four presets, and you can change them per message rather than setting one and living with it.

Calm

Steady and consistent, with very little variation between takes. The one to use when the message is information rather than persuasion, and the one that ages best across hundreds of sends.

Neutral

The balanced default. If you are not sure, this is the one, and most people never move off it once the first few land properly.

Expressive

More varied in tone, with real emotion in it. Right for a first message where the whole job is sounding like a person rather than a system.

Dramatic

Strong emotional range. Genuinely useful in small doses and conspicuous in large ones. Worth trying once on a message that is meant to land hard, and worth being honest with yourself about the result.

The practical advice, having watched people do this: start on neutral, send fifty, and only then decide you need something else. Almost everybody reaches for dramatic first and quietly moves back within a week.

Inflowave voice note composer open over a live conversation, with the message written, audio tags for emotion, reaction and pacing, and the four delivery styles calm, neutral, expressive and dramatic with expressive selected
Written, styled and sent without leaving the conversation

Recorded or generated, per note

Every note can be either an actual recording you made or speech generated from text, and the choice is per note rather than a global setting. Most people end up in the same place, and it is worth starting there.

Record the first one properly

It is the same for everybody anyway, so there is no reason to generate it, and it is the message that decides whether the person believes any of this is real. Spend ten minutes and get one you are happy with. Background noise and a slightly imperfect take help rather than hurt here.

Generate the ones that have to be specific

Once you know something about the person, a generated note can say that thing, and saying the specific thing is the entire difference between personal and personalised. A perfectly recorded generic note loses to a slightly synthetic one that mentions what they actually asked about.

Will people be able to tell?

Sometimes, and it matters less than people expect, because what gives it away is almost never the audio. It is the content. A generated note that says something only you could know about that specific person reads as real. A flawlessly rendered note saying something generic reads as a robot, and no amount of audio quality rescues it.

And if somebody asks outright whether they are speaking to a real person, answer honestly. That is not a marketing decision. It is the difference between a customer and a complaint, and it is the fastest way we have seen anybody lose an account.

Who this is built for

Creators, where your voice is the product

People followed you because of you, then message you and get a reply that could have come from anyone. A voice note closes that gap faster than any amount of good writing, because your audience already knows what you sound like. It is the one channel where being recognisable is worth more than being polished.

The volume that makes this necessary is handled in the shared inbox.

Coaches and consultants selling high ticket

The gap between somebody being interested and somebody booking a call is almost never missing information. It is that they have not decided whether they trust you, and text is a poor instrument for that. Thirty seconds of your voice, at the point they hesitated, does work that a follow-up sequence cannot.

Then the call itself is booked in the thread. See also how coaches use Inflowave.

Agencies running DMs on behalf of a client

The lead followed your client, not you, so a setter typing on their behalf is always slightly the wrong person. A clone of the client voice fixes that, and it is the case where the consent rule needs saying out loud: this is the client voice, so it is the client decision, in writing, before you record anything.

More in the agency use case and what we build for agencies.

Small and medium businesses, including local and service trades

Salons, clinics, studios, trades, dealerships. Every competitor in your area replies to enquiries with the same flat templated text, usually late. A thirty-second note from the actual owner, in the actual accent, is the cheapest differentiation available to a business this size, and it is nearly impossible for a larger competitor to copy convincingly.

The honest advice for this reader is to use less of it than you are about to. One recorded note answering the question you are asked forty times a week, and one for the objection you always get about price. That is it. A business with two good voice notes beats one with fifteen generated ones, because the fifteen start to sound like a system and the two sound like a person who could not be bothered to type.

See how small businesses use Inflowave and what it costs for a small team.

Setting it up

  1. 1. Record a sample in a quiet room

    Quality beats length by a wide margin. A few minutes on a decent microphone in a room with soft furnishings will beat half an hour in a car. Talk at the pace you actually talk rather than reading formally, because the clone learns your rhythm as much as your tone, and a stiff sample gives you a stiff voice forever.

  2. 2. Build the objection note first

    Not the introduction, which is what everybody starts with. The objection note is the one that changes outcomes, because it lands at the moment a written reply is most obviously a template. Take the objection you get most often and answer it out loud, once, properly.

  3. 3. Set the phases wider than feels right

    The instinct is to send the first note early, and the instinct is wrong. Early is exactly when it reads as automation. Let a few exchanges happen first, and you will get a better reception from fewer notes.

  4. 4. Send twenty by hand before automating any

    You will learn what lands, and more usefully what makes people go quiet, in an afternoon. Automating a script you have never said out loud is how people end up with a beautifully configured system that nobody replies to.

  5. 5. Then use fewer than you planned

    Voice works because it is unusual. A conversation with one well-placed voice note feels personal; a conversation with five feels like a broadcast, and the fifth undoes the first. Restraint is not a limitation here, it is the technique.

Against the three things people use instead

Nobody chooses between these, because they are different products. That is the point: you end up with all three and a gap in the middle. The first two rows are ours to lose and we do.

Voice capability by product category
CapabilityInflowaveVoice AI toolsManyChatRingless voicemail
Clone a voice from a sampleYesYesNoNo
Voice quality, languages and editingPartlyYesNoNo
Deliver a GENERATED voice note into an Instagram DMYesNoNoNo
Sends at the right point in the conversation, on its ownYesNoPartlyNo
The same voice answers a specific objectionYesNoPartlyNo
Send a one-off voice note by hand mid-threadYesPartlyNoNo
Reaches somebody who never gave you a phone numberYesNoYesNo
Arrives as a reply, not a broadcastYesNoPartlyNo
Delivery style you can change per messageYesYesNoNo
Use a real recording instead of a generated oneYesPartlyYesYes
The note is attached to the lead record afterwardsYesNoPartlyNo
  • Clone a voice from a sample: The dedicated voice tools do this better than us and it is not close: more voices, more languages, dubbing, editing. If the audio itself is the product, that is where to go.
  • Voice quality, languages and editing: Their entire product against one feature of ours. We are not competing here and the page says so.
  • Deliver a GENERATED voice note into an Instagram DM: The most important row on this page. ManyChat supports audio on Instagram only as a static file uploaded in Flow Builder; their documentation states that sending audio through Dynamic Block responses or the API is not supported on Instagram or WhatsApp, only Messenger and Telegram. So the voice-tool to automation-platform pipeline people try does not reach an Instagram DM at all. A voice tool on its own produces a file and stops.
  • Sends at the right point in the conversation, on its own: Notes are held per conversation phase and released on how many exchanges have actually happened, so the introduction note cannot fire on message forty. ManyChat can place a static audio block at a step in a flow, which is a different thing: it is a position in a script rather than a read of where the conversation has got to.
  • The same voice answers a specific objection: The conversation can be held when a known objection appears and a voice note assigned to that objection is used instead of a text reply.
  • Send a one-off voice note by hand mid-thread: Type it, pick a delivery style, send it into the thread you are already in. With a separate voice tool this is generate, download, upload, which is enough friction that nobody does it twice.
  • Reaches somebody who never gave you a phone number: Ringless voicemail needs a number and lands in a voicemail box. This is a reply inside a conversation the person started, which is a different permission and a different reception.
  • Arrives as a reply, not a broadcast: Ringless voicemail is outbound to a list. Whatever you think of that, it is not the same act as answering somebody who messaged you, and it carries regulatory exposure this does not.
  • Delivery style you can change per message: Four presets from steady to strongly expressive. The voice tools expose the same controls and more of them.
  • Use a real recording instead of a generated one: Any note can be an actual recording you made rather than generated speech, which is what a lot of people settle on for the first message and generation for the rest.
  • The note is attached to the lead record afterwards: It sits on the conversation and the contact like any other message, so the next person to open the deal can hear what was actually said.

Where the others are stronger. The dedicated voice tools are better than us at voice, and it is not a close call. More voices, far more languages, dubbing, fine editing, studio controls, and a team working on nothing else. If what you need is a great-sounding audio file, that is what they are for, and this page is not trying to talk you out of one. ManyChat is also a more mature DM automation product with a bigger template library and a far larger ecosystem around it, and if you only ever want to attach the same static audio clip to a step in a flow, it does that.

The gap everybody hits is not making the audio, it is getting it to the right person at the right moment without a human doing it. A voice tool hands you a file. ManyChat, by its own documentation, cannot send a generated audio file into an Instagram DM at all, only a static one placed in a flow. So the pipeline people actually try, generate the note somewhere, pass it through an automation platform, deliver it to the lead, does not exist on the channel most of them care about. Here the clone, the timing and the thread are one system: the note is held until the conversation has genuinely reached the point it was written for, and then it goes, in your voice, as a reply.

One at a time

The table is the evidence. This is the verdict, with the case for choosing them stated first, because for most people reading this the answer is genuinely one of them.

Inflowave vs a dedicated voice AI tool

Buy theirs when: Better voice, more of it, in more languages, with editing and dubbing and controls we do not have. If you are producing narration, a podcast, an audiobook or ads, buy one of those and do not think about it again. Even for DMs, if you only send a handful a week, generating a file by hand is perfectly reasonable.

Buy ours when: They finish at the file. Everything after that is you: downloading it, knowing which lead it is for, knowing whether now is the right moment, and uploading it into the thread. That works at ten conversations and collapses at two hundred, which is exactly the point at which the voice note would have been worth sending.

Inflowave vs ManyChat

Buy theirs when: A more mature DM automation product with a much bigger ecosystem, more templates and more integrations, and it does send audio on Instagram when the clip is a static file placed in a flow. If your voice note is the same clip for everybody and the flow is simple, that is a perfectly good answer and a cheaper one.

Buy ours when: Their own documentation says audio through Dynamic Block responses and the API is not supported on Instagram or WhatsApp, only on Messenger and Telegram. That rules out the whole pattern of generating a note per lead and delivering it, on the channel most people are asking about. Their docs also note the audio renders as a plain audio player rather than a native voice note. And a clip attached to a step in a flow fires by position in a script rather than by how far the conversation has actually got.

Inflowave vs ringless voicemail

Buy theirs when: If you have phone numbers and permission and your job is to reach a large list quickly, ringless voicemail does that and this does not. It is a volume tool and it is honest about being one.

Buy ours when: It is broadcast into a voicemail box, needs a phone number, and carries regulatory exposure that answering a DM does not. This is the opposite shape: one person who messaged you, getting a reply in a voice they already recognise, on the app they were already in.

Taken from vendor documentation, checked 30 July 2026. Channel support in particular changes, so verify any row that decides your choice.

Common questions

Is voice cloning illegal?

Cloning your own voice is not. Cloning the voice of another person without permission can be, and the law has moved quickly here: several US states now have specific statutes on synthetic voice and likeness, and impersonating a real person for commercial gain is actionable in most places regardless. Our own rule is simpler than the law and easier to remember: it must be your own voice. Not a celebrity, not a competitor, not a colleague who has not asked for it. If you would not be comfortable telling the person whose voice it is, do not clone it.

Can I clone my voice?

Yes, from a sample of you speaking normally. The single biggest factor in whether it sounds like you is the quality of what you record, not the length of it: a few minutes in a quiet room with a decent microphone beats half an hour of a noisy one. Record at the pace you actually talk rather than reading formally, because the clone learns your rhythm as much as your tone, and a stiff sample produces a stiff voice.

How is this different from ElevenLabs or another voice tool?

They are better at voice than we are, plainly. More voices, more languages, dubbing, editing, a whole studio. What they give you at the end is a file, and the file is not the problem. The problem is knowing which lead it is for, whether now is the right moment, and getting it into the thread. That is fine at ten conversations a week and impossible at two hundred, which is exactly when a voice note would have been worth sending.

Can I not just use ElevenLabs with ManyChat?

On Instagram, no, and this catches a lot of people out after they have built it. ManyChat supports audio on Instagram as a static file placed in a flow, but their documentation states that sending audio through Dynamic Block responses or the API is not supported on Instagram or WhatsApp, only on Messenger and Telegram. Generating a note per lead and delivering it is exactly the dynamic path, so that pipeline does not reach an Instagram DM. Their docs also note the audio plays as a standard audio player rather than a native voice note.

When does the voice note actually get sent?

When the conversation has earned it. Notes are assigned to a phase rather than to a step in a script, and a phase is defined by how many exchanges have genuinely happened. So the introduction note goes early, the warm one only once there has been a real back and forth, and the closing one later still. The point is that a voice note landing at the wrong moment is worse than no voice note, because it reads as an automation rather than a person.

Can I send one myself, mid-conversation?

Yes, and most people use it that way more than they expect to. Type what you want to say, pick a delivery style, and it goes into the thread you are already in. It is the fastest way to answer something awkward, because two sentences in your voice does what four paragraphs of text cannot. See the shared inbox.

Do I have to generate every note? Can I use real recordings?

Either. Any note can be an actual recording you made rather than generated speech, and a lot of people settle on a hybrid: record the first message properly, because it is the one that decides whether somebody believes this is real, and generate the rest so they can be specific to the person.

Will people be able to tell?

Sometimes, and the honest answer is that it depends much more on what you say than on how it sounds. A generated note saying something only you could know about that specific person reads as real. A perfectly rendered note saying something generic reads as a robot no matter how good the audio is. The technology is not the weak link here; the writing is.

Should I tell people it is AI?

It is your voice, saying something you decided to say, so most people treat it the way they treat a typed message that was drafted with help. That said, if somebody asks directly whether they are talking to a real person, answer honestly. That one is not a marketing question, it is the difference between a customer and a complaint, and it is the fastest way to lose an account we have seen.

Can it answer an objection in voice?

Yes. The conversation can be held when a known objection comes up, and a voice note written for that objection used instead of a text reply. This is the highest-value place to put a voice note, because an objection is precisely the moment where tone carries more than wording, and where a template reply is most obviously a template.

What happens to the note afterwards?

It sits on the conversation and on the contact like any other message, so whoever opens the deal next can hear what was actually said rather than guessing from a summary. That matters most when a setter hands to a closer. See the deal timeline.

What does it not do?

It is not a studio. No dubbing, no multi-language library, no fine audio editing, no sound design. If you are producing narration, a podcast or ads, use a dedicated voice tool and enjoy it. This exists to put your voice into a sales conversation at the right moment, and it is deliberately a small idea done properly rather than a large one done thinly.

Say it once. Send it at the right moment, every time.

Clone your own voice, write the note that answers the objection you always get, and let the conversation decide when it plays.

Get Started