MELVIS — Voice & Music Production

Artificial Intelligence
"Turn the sound off and a film becomes a slideshow. Turn the picture off and it is still a story. Sound was never the second half of the work."

A recording nobody has time to listen to is the same as a recording nobody made. Hours of calls, lessons and interviews sit unopened because turning them into something readable has always cost a person a day.

MELVIS listens, speaks and scores — transcription that survives a noisy room, natural speech in many languages, and original music written to a picture that already exists.

The Half of the Work That Gets Cut +

MELVIS The Half of the Work That Gets Cut

Recordings Nobody Reads

Support calls, field interviews and classroom hours pile up as audio. Somebody has to sit through all of it in real time to find the one minute that mattered, so mostly nobody does.

Voice Work Is Priced Per Language

A studio, a performer and an engineer per language means the second language usually never happens, and a product that speaks only one of them reaches only one audience.

Music Is a Licensing Problem

Cleared tracks are expensive, the cheap ones are recognisable from a hundred other campaigns, and a piece that fits the cut exactly almost never exists at the length the cut needs.

Accents and Noise Break the Tools

Transcription that works on a quiet studio voice falls apart on a market street, a phone line or an accent that was thin in the training data — which is exactly where the useful recordings are made.

Core Technologies +

MELVIS Core Technologies

Speech Recognition

Models trained across many languages and recording conditions turn speech into text with timings attached, so every sentence can be traced back to the second it was said.

Speech Synthesis

Compact text-to-speech that runs fast enough to answer a caller while they are still on the line, rather than a large model that sounds excellent and arrives too late to be an answer.

Music Generation

Models that write and perform a piece from a written description, giving a cue at the length a cut actually needs instead of one trimmed from a library track.

Long Recordings, Not Clips

An hour of audio is handled as an hour, not as a clip — split, recognised and stitched back together with the timings kept true across the joins, because the useful recordings are always the long ones.

Running It Ourselves

Open weights held on our own machines, which is what lets a hospital, a school or a call centre send us audio that was never allowed to leave their building in the first place.

What Comes Out of the Studio +

MELVIS What Comes Out of the Studio

Transcripts and Searchable Audio

Hours of recording turned into text with timings attached to every passage, so an archive that used to be listened to end to end becomes an archive that can be searched and jumped into.

Transcripts in a Reading Language

Speech recorded in one language written out in another, so a support call, a field interview or a lecture can be read by a team that does not share the language it happened in.

Voice for Systems That Answer

Spoken responses produced fast enough to arrive while a caller is still waiting, for help lines, kiosks and assistants — where a voice that comes late is the same as no voice at all.

Narration and Dubbing

Spoken versions of written material in several languages, for a course, a manual, an announcement or a film that was only ever cut in one of them.

Original Score and Cues

Music written for a specific piece at a specific length, with alternate versions for other cuts of the same campaign, cleared because we made it.

Sound for Generated Picture

Voice and music laid against film that MINEZ generated, so a piece leaves ÁRKMORA finished rather than as a silent sequence waiting on someone else.

From Recording to Finished Sound +

MELVIS From Recording to Finished Sound

Take In the Audio

Recordings arrive in whatever shape they were made, are normalised to a working format, and are logged with their source and their consent status before anything is run.

Listen and Transcribe

Speech is recognised with timings attached passage by passage, and the stretches the models were least certain about are flagged rather than quietly smoothed over.

Speak or Score

Approved text is voiced in the languages requested, or a cue is written to the picture and the timing it has to land on, and a person listens to the result.

Mix and Deliver

Levels, edit and export to the formats the client runs, with the disclosure record attached to the delivery rather than added afterwards on request.

System Architecture +

MELVIS System Architecture

Audio Layer

Source recordings held with their format, their origin and their permissions, so any transcript or voice can be traced back to the material it legitimately came from.

Model Layer

Recognition, synthesis and music models held separately behind one interface, which means a better model for one job replaces that job alone and nothing else moves.

Processing Layer

Queued jobs with settings recorded per run, because a transcript that cannot be produced again the same way is not evidence of anything.

Review Layer

Low-confidence passages and generated voice lines surfaced for a person to correct, with the correction stored against the audio that produced it.

Delivery Layer

Text, audio and music exported together with their provenance and consent records, so that what leaves the building can be accounted for later.

Whose Voice It Is +

MELVIS Whose Voice It Is

No Voice Without Consent

We do not clone a real person's voice without that person's signed permission for that specific use, and permission for one project is not permission for the next one.

Generated Speech Is Disclosed

Synthetic narration is marked as synthetic in the delivery record. A listener who needs to know whether a person said this should be able to find out.

Music We Can Account For

Cues are generated for the client and delivered with a record of how they were made, so a broadcaster asking where the music came from gets a straight answer.

Transcripts Carry Their Doubt

Where recognition was uncertain, the transcript says so instead of presenting a confident guess — which matters most in the settings where transcripts are relied upon.

What We Hold Today +

MELVIS What We Hold Today

Speech to Text

whisper-large-v3 as the general recogniser across languages, whisper-large-v3-turbo as its faster sibling for the jobs that have to come back while someone is waiting, and parakeet-tdt-0.6b-v2 for English work where turnaround matters more than breadth.

Text to Speech

Kokoro-82M for narration and responsive voice, small enough to run at conversation speed on modest hardware rather than requiring a rack of its own.

Music Generation

YuE-s1-7B-anneal-en-cot for song-length material with vocals, and Ace-Step1.5 for instrumental cues and faster iteration on a brief.

How They Run Together

All four are held on our own infrastructure behind one interface, and where a model asks that its origin be named, that notice is carried with anything we ship.

Sound for Everyone Who Has None +

MELVIS Sound for Everyone Who Has None

Inside ÁRKMORA

ZUNYUAN works in music and entertainment; ORVETH needs written knowledge read aloud where reading is hard; MINEZ generates picture that arrives silent. MELVIS is the sound under all three.

Who It Serves

Media companies and studios, schools and training providers, call centres carrying thousands of recorded conversations, and game studios needing many voices at once.

What We Have Not Done Yet

We have not published word error rates, latency or cost per hour from our own bench. Those runs have not been done and recorded, so this page describes what the studio is built to do, not how well it scores.

What We Do Not Hold Yet

Telling one speaker from another in a recording of several is not something we hold a model for today. A transcript from us carries timings, not names, and we say so rather than letting a customer assume otherwise.

Where It Goes Next

Separating speakers in a crowded room, more languages that current tools handle poorly, voice that keeps one performer consistent across a long series, and music that follows an edit when the edit changes.

"Everything worth saying was said out loud first. We built the part that listens."

MELVIS — Voice & Music Production.