Küssnacht am Rigi · Canton Schwyz

The face is not mine. The voice is.

The video that opens this site's home page is built exactly that way: a computer-generated avatar, carrying Enrieta Fontana's own voice, cloned from a twenty-six second recording.

It is not a demo assembled for this page. It is the first thing anyone landing on enrietabiz.com sees — and the video says so itself, out loud.

What you see on the home page is the service

The home page presentation video. Generated face, real cloned voice. It starts muted: the sound button turns the audio on.

The voice was recorded once: twenty-six seconds in a quiet room. After that there is nothing left to record — you write the text and the voice reads it, with the same pauses and the same inflection.

The face, on the other hand, does not exist. It is a generated portrait, animated on the speech: mouth, eyes and head follow the audio. It is the part people notice first, and it is also the part that has to be declared first.

“The face you saw at the start is an avatar. And the voice is mine, cloned — I'm telling you because that is exactly the work I do.”

The same cut exists in Italian, English and German. There is still only one original recording: the other two languages were never actually spoken.

The tribute to Moira Orfei

Tribute to Moira Orfei — 72 seconds, generated portrait animated on the speech. Made for the Moira Orfei Circus Academy.

Commissioned by the Moira Orfei Circus Academy to remember its founder, who died in 2015. Seventy-two seconds: a portrait rebuilt from archive material, animated on the speech, with a voice built for the occasion.

It was the longest piece of this kind we have made, and the most delicate: here you are not inventing a character, you are touching the memory a family and an audience hold of a real person. Every step was seen and approved by the people who commissioned it before anything was edited together.

“I can only express my deepest gratitude for the beautiful work done on the video dedicated to Moira Orfei. It was a long and demanding project, a spine-tingling piece of work that brought her back to life through images with extraordinary intensity, and the final result exceeded every expectation.”
Brigitta Boccoli — actress, founder of the Moira Orfei Circus Academy

How it is actually done

Three separate steps, and you can buy them one at a time: the voice without the avatar, or the avatar with an off-the-shelf voice.

1

The voice

Thirty seconds of clean recording is enough — a phone in a room without echo will do. From then on the voice is reusable: you write the text and it reads it, as many times as you need.

2

The face

Either a real photograph of you, or a generated portrait if you would rather not put your face on it. Front-on, flat light, shoulders in frame: that is the framing that holds the animation best.

3

The video

You write the script, the voice reads it, the face speaks it. Subtitles included, imposed from the written text rather than transcribed by ear — so proper names do not come out mangled.

Three things we do not do

This technology becomes a problem the day it is used to make someone say something they never said. These are the house rules, and they hold even when a client asks for the opposite.

  • We do not clone anyone's voice or face without their written consent — or, if the person has died, without the consent of whoever has standing.
  • We do not pass a generated portrait off as a photograph. If the face does not exist, the text alongside it says so.
  • We do not put statements, opinions or prices into anyone's mouth that they have not approved line by line.

What it costs

The voice is cloned once and then it stays: it is the smallest of the three costs, and it pays for itself from the second video onwards.

The real price depends on how long the video runs, how many languages you need and how many rounds of revision you are budgeting for. A one-minute presentation video is one thing; seventy-two seconds rebuilt from archive material, like the tribute, is quite another.

We start from the content, not from a price list: tell us what the video has to say and who to, and we will tell you whether this is the right way to make it.

The rest of the work

Questions

How much audio does it take to clone a voice?

For EnrietaBiz, thirty seconds of clean recording is enough: a room without echo and a phone microphone. No studio required. The quality of the result depends more on background silence than on length — two minutes recorded badly are worth less than thirty seconds recorded well.

Can I have an avatar without using my own face?

Yes. EnrietaBiz generates a portrait that matches no existing person and animates it on the speech. That is what Enrieta Fontana did for the presentation video on enrietabiz.com: the face is generated, the voice is hers. In that case the text has to say the face is an avatar — that is a house rule, not an option.

Can the same voice speak another language?

Yes. The EnrietaBiz presentation video exists in Italian, English and German from a single Italian recording: the other two languages were never spoken. The accent stays that of the person who recorded, and in some languages you can hear it — so it gets tested before it gets decided.

Can you make a video with someone who has died?

It can be done, and EnrietaBiz has done it: the tribute to Moira Orfei, seventy-two seconds, commissioned by the Moira Orfei Circus Academy. It is only done at the request of whoever has standing — the family, the heir, the foundation — and with line-by-line approval of what the video says. Without that, it does not start.

How long does it take?

For a one-minute presentation video with a cloned voice and an avatar, EnrietaBiz works in days: the voice is cloned in an afternoon, the rest is script and revisions. Reconstruction work from archive material, like the Moira Orfei tribute, takes weeks and is quoted separately.

If you need a face that speaks

Bring the text it should say, or just the idea. Half an hour is enough to tell whether this is the right road — and if it is not, we will say so.

Book half an hour