Handbook
Drag an audio file onto the timeline, or press Ctrl+I. It lands where you drop it. Space plays and pauses. The mouse wheel zooms; Shift+F fits the whole session on screen.
Seston has two views, switched at the top left. Multitrack is where clips are arranged against each other. Waveform is one clip on its own, for close work — double-clicking a clip takes you there.
A session — a .seston file — holds the arrangement: which clips sit where, their
gains, fades, volume lines and effects. It does not hold the audio. The files you imported stay
where they are and are never written to.
The Files panel lists everything the session is using. The + button drops a file onto the selected track at the playhead; dragging a row onto the timeline puts it where you let go. Right-clicking a row offers rename, selecting every clip that uses it, and removing it from the session — which takes its clips out too, and never touches the file on disk.
Seston writes recovery data every few minutes while you work, separately from your session file,
and offers it back the next time it starts. Saying no leaves it in place; it is not thrown away
until you save properly. Saving also keeps the previous few versions of the file beside it, as
Show.1.seston and so on.
Every clip carries one, drawn flat until you use it. Double-click the line to add a handle, then drag handles to shape the level across the clip. Right-click a handle to remove it. The line belongs to the clip, so it moves and copies with it.
Snap pulls edges to the grid, to other clips' edges and to the playhead. Its reach is capped in milliseconds rather than pixels, so zooming out does not quietly turn a 50-millisecond snap into a one-second one.
For a long unedited take. It finds the dead air, takes it out, and lays the surviving phrases head-to-tail across two tracks, alternating, with every join overlapping its neighbour by a crossfade. The alternation is what makes the crossfades possible: butted cuts on a single track have nothing to fade to.
Seston will drop a held "eee" if it is short, standing alone between real pauses, quieter than the speech around it and spectrally almost motionless. That is a signal test, not a language one: a filler buried inside a sentence is deliberately left for you.
Select two or more clips and press Ctrl+G, or use Clip › Group. From then on they are selected as one and move as one. Ctrl+Shift+G frees them again.
This is what a checkerboarded edit wants once the joins are right: dragging one piece of it and leaving the rest behind is only ever a mistake. Clicking any member selects the whole group, a rubber band that catches one catches all of them, and splitting a grouped clip leaves both halves in the group — the join is exactly what the group was protecting.
A grouped clip carries a small chain mark in its header. Copying a group and pasting it gives the copies a group of their own, so moving the paste leaves the original where it is.
Once an edit is right, select the pieces and press Ctrl+M. They are rendered down to a single clip with everything they carried printed in — gains, fades, volume handles and clip effects. Gaps become silence, and clips that overlap are summed, so the crossfades in a checkerboarded edit survive the merge rather than being thrown away.
Select a clip and press Enhance. It measures the recording and works through it in one pass: rumble and DC offset out, noise and room out, the microphone's tone corrected toward a speech curve, the pauses shut, the level evened, the loudness set, the peaks limited.
The report at the end says what it measured and what it decided: the room's decay time, how far the worst band was out of place, where the pauses were shut, the gain applied. If it says the level had to be held back, the cleaning took out more than it should have — listen to it before you use it, and try again with the model off or the strength lower.
Both the recording and the cleaned version are kept, so the mix between them can be moved at any time from Clip → Enhancement mix. That is immediate however long the clip is: nothing is measured or cleaned again, the two are simply blended differently. At 0% the clip is the recording again, sample for sample.
Running Enhance a second time on the same clip also reads the original recording rather than the last result, so the two passes never compound. Cleaning already-cleaned audio takes out what is left of the room and a good deal of the voice with it, and that does not undo.
Right-click a track and choose Generate speech at the playhead, or use Speech › Generate Speech. Type or paste the line, pick a voice, and the audio arrives as an ordinary clip — cut it, enhance it, merge it like anything else.
The script is kept on the clip, so Speak this line again is there when the script needs a correction. You do not have to type it back in.
The speech is generated by ElevenLabs using your API key. Seston has no account there and no part in the billing — the usage is yours. Get a key from your ElevenLabs account settings and enter it once.
The key is encrypted with Windows' own data protection under your user account, so it cannot
be read from another account or carried to another machine. It is not in
settings.json, it is never written to a session file or the status line, and it
is never shown again after you enter it. Set key replaces it and
Remove the key forgets it.
When you want to control it yourself, double-click the clip to open it in the waveform view, then:
The order matters: the profile has to come from silence. Taking it from speech and subtracting it removes the speech.
If the meter is not moving, nothing is being captured. That is the fastest thing to check before recording anything you cannot repeat.
Set the source to System playback. Seston then records whatever the machine is playing — the far end of a phone, WhatsApp or Zoom call — without the call application being involved at all.
Loopback only hears what the machine plays, which on a call is them and not you. So your microphone is recorded at the same time onto its own track, and the two sides can be edited apart. This is on by default; without it half the conversation would go unrecorded.
Recording a conversation needs the other person's agreement in many places. Ask first.
High-pass, low-pass, EQ, gate, compressor, de-esser and limiter. They can sit on a clip (Clip › Clip Effects), on a track (the FX button in its header) or on the master (Audio › Master Effects). The rack is a window, so you can leave it open while you listen.
A track header showing FX 2 has two effects running. Eleven chains for voice come built in, and anything you set up can be saved beside them.
Clip source → clip gain and fades → clip effects → summed into the track → track effects → pan and fader → summed into the master → master effects → master gain → safety limiter.
Every channel has a meter down the right of its strip, and the master has the wide one across the transport. The bar is the average level; the line at its tip is the peak, held for a moment so a transient too short to move the bar still leaves a reading. The gap between them is the crest factor — on speech, normally 12 to 18 dB.
A red mark at the end means something reached full scale. Click a meter to clear it.
WAV at 16, 24 or 32-bit float, FLAC, MP3 from 64 to 320 kbps, or AAC. Sample rates from 22.05 up to 192 kHz, stereo or mono. The MP3, AAC and FLAC encoders are the ones Windows ships, so nothing extra is bundled and nothing is licensed on your behalf.
Set Loudness in the export dialog and Seston measures the whole mix to ITU-R BS.1770 first, then delivers it at the target with the peaks limited to the ceiling rather than clipped. −16 LUFS is podcast, −14 is what the streaming services normalise to, −19 is radio playout, −23 is EBU R128 broadcast.
Measuring means rendering twice, so it takes about twice as long. That is the whole point: loudness is a property of the whole programme and there is no honest way to know it from the first few seconds. The report at the end is measured in the file that was written — if it says −16.0 LUFS, that is what the file is.
Mono halves the file and loses nothing a single voice put there. Note that a stereo mix folded to mono is not the loudness the stereo was, so the measurement is taken after the fold, not before.
The combinations that come up are named: podcast, streaming, radio playout, broadcast, archive, voice note, master. Picking one sets the format, rate, channels and loudness together; changing any of them afterwards is allowed and moves the preset to Custom.
Edit › Preferences, or Ctrl+,.
Ctrl+, and look at Audio Hardware. The page shows which device is actually open, which is the only way to tell whether a WASAPI request quietly fell back to MME. If in doubt, choose MME and the system default.
Watch the meter while recording — it shows the incoming signal, not the mix. If it is not moving, the source is wrong. On System playback, remember it captures what the machine plays, so something has to be playing.
Seston has not been through code signing. Choose More info, then Run anyway.
Try Clip → Enhancement mix first and bring it down to 50 or 60% — that is immediate, and most of the time it is what a voice that already sounded reasonable wanted. If it is still wrong, run Enhance again with the speech model switched off or the strength lower; it reads the original recording, so a second pass does not add to the first. If the report said the level had to be held back, that is the sign the cleaning took out too much.