
"Turn the sound off and a film becomes a slideshow. Turn the picture off and it is still a story. Sound was never the second half of the work."
A recording nobody has time to listen to is the same as a recording nobody made. Hours of calls, lessons and interviews sit unopened because turning them into something readable has always cost a person a day.
MELVIS listens, speaks and scores — transcription that survives a noisy room, natural speech in many languages, and original music written to a picture that already exists.

Support calls, field interviews and classroom hours pile up as audio. Somebody has to sit through all of it in real time to find the one minute that mattered, so mostly nobody does.
A studio, a performer and an engineer per language means the second language usually never happens, and a product that speaks only one of them reaches only one audience.
Cleared tracks are expensive, the cheap ones are recognisable from a hundred other campaigns, and a piece that fits the cut exactly almost never exists at the length the cut needs.
Transcription that works on a quiet studio voice falls apart on a market street, a phone line or an accent that was thin in the training data — which is exactly where the useful recordings are made.

Models trained across many languages and recording conditions turn speech into text with timings attached, so every sentence can be traced back to the second it was said.
Compact text-to-speech that runs fast enough to answer a caller while they are still on the line, rather than a large model that sounds excellent and arrives too late to be an answer.
Models that write and perform a piece from a written description, giving a cue at the length a cut actually needs instead of one trimmed from a library track.
An hour of audio is handled as an hour, not as a clip — split, recognised and stitched back together with the timings kept true across the joins, because the useful recordings are always the long ones.
Open weights held on our own machines, which is what lets a hospital, a school or a call centre send us audio that was never allowed to leave their building in the first place.

Hours of recording turned into text with timings attached to every passage, so an archive that used to be listened to end to end becomes an archive that can be searched and jumped into.
Speech recorded in one language written out in another, so a support call, a field interview or a lecture can be read by a team that does not share the language it happened in.
Spoken responses produced fast enough to arrive while a caller is still waiting, for help lines, kiosks and assistants — where a voice that comes late is the same as no voice at all.
Spoken versions of written material in several languages, for a course, a manual, an announcement or a film that was only ever cut in one of them.
Music written for a specific piece at a specific length, with alternate versions for other cuts of the same campaign, cleared because we made it.
Voice and music laid against film that MINEZ generated, so a piece leaves ÁRKMORA finished rather than as a silent sequence waiting on someone else.

Recordings arrive in whatever shape they were made, are normalised to a working format, and are logged with their source and their consent status before anything is run.
Speech is recognised with timings attached passage by passage, and the stretches the models were least certain about are flagged rather than quietly smoothed over.
Approved text is voiced in the languages requested, or a cue is written to the picture and the timing it has to land on, and a person listens to the result.
Levels, edit and export to the formats the client runs, with the disclosure record attached to the delivery rather than added afterwards on request.

Source recordings held with their format, their origin and their permissions, so any transcript or voice can be traced back to the material it legitimately came from.
Recognition, synthesis and music models held separately behind one interface, which means a better model for one job replaces that job alone and nothing else moves.
Queued jobs with settings recorded per run, because a transcript that cannot be produced again the same way is not evidence of anything.
Low-confidence passages and generated voice lines surfaced for a person to correct, with the correction stored against the audio that produced it.
Text, audio and music exported together with their provenance and consent records, so that what leaves the building can be accounted for later.

We do not clone a real person's voice without that person's signed permission for that specific use, and permission for one project is not permission for the next one.
Synthetic narration is marked as synthetic in the delivery record. A listener who needs to know whether a person said this should be able to find out.
Cues are generated for the client and delivered with a record of how they were made, so a broadcaster asking where the music came from gets a straight answer.
Where recognition was uncertain, the transcript says so instead of presenting a confident guess — which matters most in the settings where transcripts are relied upon.

whisper-large-v3 as the general recogniser across languages, whisper-large-v3-turbo as its faster sibling for the jobs that have to come back while someone is waiting, and parakeet-tdt-0.6b-v2 for English work where turnaround matters more than breadth.
Kokoro-82M for narration and responsive voice, small enough to run at conversation speed on modest hardware rather than requiring a rack of its own.
YuE-s1-7B-anneal-en-cot for song-length material with vocals, and Ace-Step1.5 for instrumental cues and faster iteration on a brief.
All four are held on our own infrastructure behind one interface, and where a model asks that its origin be named, that notice is carried with anything we ship.

ZUNYUAN works in music and entertainment; ORVETH needs written knowledge read aloud where reading is hard; MINEZ generates picture that arrives silent. MELVIS is the sound under all three.
Media companies and studios, schools and training providers, call centres carrying thousands of recorded conversations, and game studios needing many voices at once.
We have not published word error rates, latency or cost per hour from our own bench. Those runs have not been done and recorded, so this page describes what the studio is built to do, not how well it scores.
Telling one speaker from another in a recording of several is not something we hold a model for today. A transcript from us carries timings, not names, and we say so rather than letting a customer assume otherwise.
Separating speakers in a crowded room, more languages that current tools handle poorly, voice that keeps one performer consistent across a long series, and music that follows an edit when the edit changes.
"Everything worth saying was said out loud first. We built the part that listens."
MELVIS — Voice & Music Production.