The voice
Thirty seconds of clean recording is enough — a phone in a room without echo will do. From then on the voice is reusable: you write the text and it reads it, as many times as you need.
The video that opens this site's home page is built exactly that way: a computer-generated avatar, carrying Enrieta Fontana's own voice, cloned from a twenty-six second recording.
It is not a demo assembled for this page. It is the first thing anyone landing on enrietabiz.com sees — and the video says so itself, out loud.
The home page presentation video. Generated face, real cloned voice. It starts muted: the sound button turns the audio on.
The voice was recorded once: twenty-six seconds in a quiet room. After that there is nothing left to record — you write the text and the voice reads it, with the same pauses and the same inflection.
The face, on the other hand, does not exist. It is a generated portrait, animated on the speech: mouth, eyes and head follow the audio. It is the part people notice first, and it is also the part that has to be declared first.
“The face you saw at the start is an avatar. And the voice is mine, cloned — I'm telling you because that is exactly the work I do.”
The same cut exists in Italian, English and German. There is still only one original recording: the other two languages were never actually spoken.
Commissioned by the Moira Orfei Circus Academy to remember its founder, who died in 2015. Seventy-two seconds: a portrait rebuilt from archive material, animated on the speech, with a voice built for the occasion.
It was the longest piece of this kind we have made, and the most delicate: here you are not inventing a character, you are touching the memory a family and an audience hold of a real person. Every step was seen and approved by the people who commissioned it before anything was edited together.
“I can only express my deepest gratitude for the beautiful work done on the video dedicated to Moira Orfei. It was a long and demanding project, a spine-tingling piece of work that brought her back to life through images with extraordinary intensity, and the final result exceeded every expectation.”
Three separate steps, and you can buy them one at a time: the voice without the avatar, or the avatar with an off-the-shelf voice.
Thirty seconds of clean recording is enough — a phone in a room without echo will do. From then on the voice is reusable: you write the text and it reads it, as many times as you need.
Either a real photograph of you, or a generated portrait if you would rather not put your face on it. Front-on, flat light, shoulders in frame: that is the framing that holds the animation best.
You write the script, the voice reads it, the face speaks it. Subtitles included, imposed from the written text rather than transcribed by ear — so proper names do not come out mangled.
This technology becomes a problem the day it is used to make someone say something they never said. These are the house rules, and they hold even when a client asks for the opposite.
The voice is cloned once and then it stays: it is the smallest of the three costs, and it pays for itself from the second video onwards.
The real price depends on how long the video runs, how many languages you need and how many rounds of revision you are budgeting for. A one-minute presentation video is one thing; seventy-two seconds rebuilt from archive material, like the tribute, is quite another.
We start from the content, not from a price list: tell us what the video has to say and who to, and we will tell you whether this is the right way to make it.
The agents that answer your clients for you
Custom AI agents →The flows that take the repetitive work away
n8n automation →What the people we have worked for say
Reviews →For EnrietaBiz, thirty seconds of clean recording is enough: a room without echo and a phone microphone. No studio required. The quality of the result depends more on background silence than on length — two minutes recorded badly are worth less than thirty seconds recorded well.
Yes. EnrietaBiz generates a portrait that matches no existing person and animates it on the speech. That is what Enrieta Fontana did for the presentation video on enrietabiz.com: the face is generated, the voice is hers. In that case the text has to say the face is an avatar — that is a house rule, not an option.
Yes. The EnrietaBiz presentation video exists in Italian, English and German from a single Italian recording: the other two languages were never spoken. The accent stays that of the person who recorded, and in some languages you can hear it — so it gets tested before it gets decided.
It can be done, and EnrietaBiz has done it: the tribute to Moira Orfei, seventy-two seconds, commissioned by the Moira Orfei Circus Academy. It is only done at the request of whoever has standing — the family, the heir, the foundation — and with line-by-line approval of what the video says. Without that, it does not start.
For a one-minute presentation video with a cloned voice and an avatar, EnrietaBiz works in days: the voice is cloned in an afternoon, the rest is script and revisions. Reconstruction work from archive material, like the Moira Orfei tribute, takes weeks and is quoted separately.
Bring the text it should say, or just the idea. Half an hour is enough to tell whether this is the right road — and if it is not, we will say so.
Book half an hour