Vardast

Voice Messages in Your DMs: What AI Hears and What It Sells

Category: DM Automation
Voice Messages in Your DMs: What AI Hears and What It Sells

Customers keep sending voice notes, and sellers keep leaving them unplayed. Here is what AI transcription really does with those recordings, where it breaks, and how to turn audio into orders.

It is 11:40 on a Tuesday night. A seller's inbox has fourteen unread messages. Thirteen are text. One is a voice note, forty-seven seconds long, with that little unplayed dot sitting next to it.

Guess which one she opened last.

She got to it three days later. By then the customer had bought the same bag somewhere else. The recording, when it finally played, said: "Hi, do you still have the black one in medium, do you ship to my city, and can I pay on delivery?" Three buying questions in one file. Zero answers.

Voice notes are the most under-answered message type in DM selling, and it is not because sellers are lazy. A voice note cannot be skimmed. Text you read in two seconds while walking to the car. Audio makes you stop, find a quiet room, and hand over forty-seven seconds you do not have.

Why your customers keep pressing record

Nobody sends a voice note to make your life harder. They send one because typing is the expensive part of their day.

Writing a paragraph on a phone keyboard takes a minute. Saying the same thing takes ten seconds. If your buyer is on their feet, driving, holding a kid, or fighting an autocorrect that does not speak their language properly, recording wins every single time.

There is an age pattern too. Older buyers send more audio, and they send longer audio. In markets where phones arrived before laptops, voice is not a quirk. It is the normal way a person talks to a shop.

Here is the part most sellers miss: a voice note is a signal of effort. Nobody records forty seconds about a product they are mildly curious about. Someone who speaks into your inbox has already half-decided you might be where they buy. That unplayed audio is usually warmer than the "hi" sitting above it, and it is rotting while you sleep. DM buyers wait far less than sellers think they do.

What actually happens between "record" and "reply"

The mechanics are simple enough. The audio file arrives, an AI assistant converts speech to text, and the text becomes a normal message in the conversation. Transcription is the easy half.

The hard half is what the words do next, because spoken language is a mess. No punctuation. False starts. Two topics glued together with an "and also". A real transcript reads closer to this:

"Hi, sorry, the green one, the one you posted in the story yesterday, is it still there, and how much was it, I think you said eight ninety but I am not sure, and does it come in a bigger size because it is for my sister actually..."

A keyword bot looks at that and picks one word. Usually the wrong one. An assistant worth paying for treats it as a paragraph of intent: it separates the four questions, answers them in order, and asks about the one thing that was genuinely unclear. That is the whole difference. Leaving even one of those questions hanging is expensive, and I have written before about how a few unanswered questions quietly kill sales.

The three voice notes you will get this week

After you watch enough seller inboxes, the shapes repeat.

  • The stacked question. Thirty to sixty seconds, three or four questions, one breath. Highest intent of anything in your inbox. Also the one most likely to get a reply that answers only the last question asked.
  • The half-sentence. Four seconds: "Hi, how much is it?" No product named, no context, sent after tapping a story. The answer is not a price. The answer is a question back, plus the story context you already have.
  • The complaint. The longest recordings you will ever receive, and the most emotional. People record when they are upset because typing anger takes too long. These need a human faster than anything else in the inbox.

Those three want three different handoff rules. If you set up only one, set up the third.

Where voice breaks, and how to make it break less

Transcription is good now. It is not magic, and the places it slips are boringly predictable:

  • Product names. Brand names, model codes and the nickname your customers invented for a product get mangled first. Your assistant needs those spellings and nicknames written into its product knowledge, including the wrong ones people say out loud.
  • Numbers. Sizes, prices, order codes and phone numbers are where a small error becomes a refund. Spoken numbers slur.
  • Noise. Traffic, a TV, a shop counter. Some audio simply will not resolve cleanly, and pretending otherwise is worse than admitting it.
  • Two people at once. Someone recording while asking their partner what colour they wanted. The transcript will contain both voices with no labels.

So build one habit into your assistant: never let it guess. When a transcript is only partly clear, the reply should repeat back what it understood and ask. "I understood you are asking about the black bag in medium, with delivery to your city, is that right?" One extra message. Compare that to a wrong shipment, a return, a refund and a follower who tells three friends.

And always echo numbers back in writing. If someone says a price out loud, your reply should show it as digits. That single rule prevents most of the damage transcription can do.

Reply in text. Almost always.

Sellers who love voice notes often want to answer with one. I think that is a mistake, and I will defend it.

Your customer cannot skim your voice note either. They cannot screenshot the price and send it to their sister. They cannot search "size" three weeks later when they want to reorder. They cannot copy your card number or your address. If you record a ninety-second answer, you have handed your own problem straight back to the person who was about to pay you.

There is one exception worth making. An angry customer, mid-complaint, sometimes needs to hear an actual human being sound sorry. Fine. Send the short human voice reply. Then send the price, the address and the order number as text right after, because that is the part they need to read twice.

The real test is the second message

Handling one voice note is a demo. Handling the follow-up is a business.

What usually happens is this: the customer records a question, you answer, and then they type "and the blue one?" Two words. No context. Different message type, same conversation. If your assistant treated the voice note as a one-off ticket, it now has no idea what "the blue one" refers to, and the customer has to start over. Nobody starts over. They leave.

Whatever you use has to carry the audio, the transcript and the text into one thread with one memory, across every channel the customer might reappear on. That is exactly the argument for one inbox with one memory rather than four apps that each know a third of the story.

What to set up this week

  1. Open every unplayed voice note older than a day. Not to reply. To read what your customers were actually asking, in their words.
  2. Write the five most common of those questions into your assistant's knowledge, with the answers you would give at your best, not your most tired.
  3. Add the nicknames and misspellings people say out loud for your top products.
  4. Set the rule: unclear audio gets a confirming question, never a guess.
  5. Route any recording that sounds like a complaint to a human immediately, day or night.

Do those five and the forty-seven-second file stops being a chore you postpone.

This is one of the reasons Vardast listens to voice messages instead of ignoring them: the recording gets understood, every question inside it gets answered in your brand voice from your own product knowledge, and anything emotional or unclear goes to you with the transcript already attached. The seller does not lose the sale to a blue dot.

The customer who records is the one who already wants to buy. Answering that within minutes is the cheapest growth available to you, and it costs nothing but the decision to stop letting audio pile up.

Voice Messages and AI Transcription: Questions Sellers Ask

Can an AI assistant really understand voice messages in Instagram or WhatsApp DMs?

Yes. The audio is converted to text and then handled like any other message, so the assistant can answer from your product knowledge. Quality depends on the recording: clear speech works well, heavy background noise and shouted numbers are where errors appear.

Do I need a separate transcription tool for customer audio?

Not if your assistant already handles voice as a native message type. Bolting on a standalone transcription service usually means you get a text file with no conversation memory around it, which leaves you doing the reply work anyway.

Should my assistant reply with a voice message too?

Usually no. Customers cannot skim, screenshot or search audio, so prices, addresses and order details belong in text. The one reasonable exception is a short human voice reply to an upset customer, followed immediately by the details in writing.

What happens with strong accents or dialects?

Common accents are handled well and unusual dialect words are where transcripts get thin. The safe design is a confirming question when something is unclear, so a half-heard product name never turns into a wrong shipment.

Are voice notes actually worth answering, or are they mostly time-wasters?

They are among the highest intent messages you receive. Recording forty seconds takes effort that casual browsers do not spend, so the unplayed audio in your inbox is typically warmer than the short text messages sitting above it.