Handbook

Getting started

Drag an audio file onto the timeline, or press Ctrl+I. It lands where you drop it. Space plays and pauses. The mouse wheel zooms; Shift+F fits the whole session on screen.

Seston has two views, switched at the top left. Multitrack is where clips are arranged against each other. Waveform is one clip on its own, for close work — double-clicking a clip takes you there.

Sessions and files

A session — a .seston file — holds the arrangement: which clips sit where, their gains, fades, volume lines and effects. It does not hold the audio. The files you imported stay where they are and are never written to.

The Files panel lists everything the session is using. The + button drops a file onto the selected track at the playhead; dragging a row onto the timeline puts it where you let go. Right-clicking a row offers rename, selecting every clip that uses it, and removing it from the session — which takes its clips out too, and never touches the file on disk.

If something goes wrong

Seston writes recovery data every few minutes while you work, separately from your session file, and offers it back the next time it starts. Saying no leaves it in place; it is not thrown away until you save properly. Saving also keeps the previous few versions of the file beside it, as Show.1.seston and so on.

Editing clips

  • Drag the body of a clip to move it. Drag it up or down to change track.
  • Drag either end to trim. This changes the clip's window onto the file, not the file.
  • Drag the top corners for a fade in or out.
  • Alt and drag leaves a copy behind.
  • S splits at the playhead. The razor tool cuts exactly where you click, with no snapping — a cut should land where you aimed it.

Volume lines

Every clip carries one, drawn flat until you use it. Double-click the line to add a handle, then drag handles to shape the level across the clip. Right-click a handle to remove it. The line belongs to the clip, so it moves and copies with it.

Snapping

Snap pulls edges to the grid, to other clips' edges and to the playhead. Its reach is capped in milliseconds rather than pixels, so zooming out does not quietly turn a 50-millisecond snap into a one-second one.

Smart Edit

For a long unedited take. It finds the dead air, takes it out, and lays the surviving phrases head-to-tail across two tracks, alternating, with every join overlapping its neighbour by a crossfade. The alternation is what makes the crossfades possible: butted cuts on a single track have nothing to fade to.

  1. Select the clip and press Smart Edit.
  2. Set how long a pause has to be before it counts as dead air — 0.45 seconds by default. Shorter pauses are part of the delivery and stay.
  3. It analyses, then reports what it found: how much shorter the result is, how many phrases, how much silence, how many held fillers.
  4. Apply, or cancel. One Ctrl+Z puts it all back.

About filler removal

Seston will drop a held "eee" if it is short, standing alone between real pauses, quieter than the speech around it and spectrally almost motionless. That is a signal test, not a language one: a filler buried inside a sentence is deliberately left for you.

Grouping clips

Select two or more clips and press Ctrl+G, or use Clip › Group. From then on they are selected as one and move as one. Ctrl+Shift+G frees them again.

This is what a checkerboarded edit wants once the joins are right: dragging one piece of it and leaving the rest behind is only ever a mistake. Clicking any member selects the whole group, a rubber band that catches one catches all of them, and splitting a grouped clip leaves both halves in the group — the join is exactly what the group was protecting.

A grouped clip carries a small chain mark in its header. Copying a group and pasting it gives the copies a group of their own, so moving the paste leaves the original where it is.

Merging clips

Once an edit is right, select the pieces and press Ctrl+M. They are rendered down to a single clip with everything they carried printed in — gains, fades, volume handles and clip effects. Gaps become silence, and clips that overlap are summed, so the crossfades in a checkerboarded edit survive the merge rather than being thrown away.

Enhance Speech

Select a clip and press Enhance. It measures the recording and works through it in one pass: rumble and DC offset out, noise and room out, the microphone's tone corrected toward a speech curve, the pauses shut, the level evened, the loudness set, the peaks limited.

The settings worth knowing

  • Use the speech model — a trained model does the noise and room. Slower than the measured stages, and better at both. Leave it on unless it is doing something you do not like.
  • Loudness target — −16 LUFS for podcast, −14 for streaming services, −23 for broadcast.
  • Strength — how hard the stages that can be heard working are pushed. Past about 0.85 a voice starts to sound processed rather than clean.
  • Mix — how much of the cleaned version ends up in the clip, against the recording it was made from. 100% is all of it. A voice that was nearly clean to begin with usually wants less.

The report at the end says what it measured and what it decided: the room's decay time, how far the worst band was out of place, where the pauses were shut, the gain applied. If it says the level had to be held back, the cleaning took out more than it should have — listen to it before you use it, and try again with the model off or the strength lower.

Changing your mind afterwards

Both the recording and the cleaned version are kept, so the mix between them can be moved at any time from Clip → Enhancement mix. That is immediate however long the clip is: nothing is measured or cleaned again, the two are simply blended differently. At 0% the clip is the recording again, sample for sample.

Running Enhance a second time on the same clip also reads the original recording rather than the last result, so the two passes never compound. Cleaning already-cleaned audio takes out what is left of the room and a good deal of the voice with it, and that does not undo.

What it cannot do. It removes what should not be there and corrects what is measurably wrong. It cannot put back detail the microphone never captured.
Sessions get bigger. A session holding an enhanced clip stores both the recording and the cleaned version, because that is what makes the mix movable. The blend itself is not stored — it is worked out again when the session opens.

Generating a voice

Right-click a track and choose Generate speech at the playhead, or use Speech › Generate Speech. Type or paste the line, pick a voice, and the audio arrives as an ordinary clip — cut it, enhance it, merge it like anything else.

  • Model — fetched from your account rather than written into Seston, so a model added next month is there without an update. Each one lists what it does and how many languages it speaks.
  • Language — leaving it on Detect from the text is right nearly always, and wrong on the sentence with an English brand name in the middle of a Turkish one. Setting it holds the reading to that language whatever the text looks like.
  • Stability — low lets the reading wander and perform; high reads the same way every time, which is what a bulletin wants and a drama does not.
  • Similarity — how closely it holds to the original voice.
  • Style — acts more. Costs time, and overdoes it near the top.

The script is kept on the clip, so Speak this line again is there when the script needs a correction. You do not have to type it back in.

Your key, your account

The speech is generated by ElevenLabs using your API key. Seston has no account there and no part in the billing — the usage is yours. Get a key from your ElevenLabs account settings and enter it once.

What leaves the machine. Seston works entirely on your own computer. This one feature is the exception: while it is used, the script and your key are sent to ElevenLabs and nowhere else. Nothing else ever leaves — not your audio, not your sessions, not what you do with them.

The key is encrypted with Windows' own data protection under your user account, so it cannot be read from another account or carried to another machine. It is not in settings.json, it is never written to a session file or the status line, and it is never shown again after you enter it. Set key replaces it and Remove the key forgets it.

Taking noise out by hand

When you want to control it yourself, double-click the clip to open it in the waveform view, then:

  1. Select a pause — a second or two with nothing but the room on it.
  2. Capture noise. That measures the shape of what is there.
  3. Reduce noise, and set how far down to push it.

The order matters: the profile has to come from silence. Taking it from speech and subtracting it removes the speech.

Recording

  1. Choose what to record from: Audio › Record From, or Preferences.
  2. Click the track it should land on, or arm one with R.
  3. Press the record button. The transport rolls and the meter shows what is coming in.
  4. Stop ends the take.

If the meter is not moving, nothing is being captured. That is the fastest thing to check before recording anything you cannot repeat.

Recording a call

Set the source to System playback. Seston then records whatever the machine is playing — the far end of a phone, WhatsApp or Zoom call — without the call application being involved at all.

Loopback only hears what the machine plays, which on a call is them and not you. So your microphone is recorded at the same time onto its own track, and the two sides can be edited apart. This is on by default; without it half the conversation would go unrecorded.

Wear headphones. On speakers the far end comes back into your microphone: they hear themselves, and the same voice ends up on both tracks a few milliseconds apart, which is worse to fix than either problem alone.

Recording a conversation needs the other person's agreement in many places. Ask first.

Effects and presets

High-pass, low-pass, EQ, gate, compressor, de-esser and limiter. They can sit on a clip (Clip › Clip Effects), on a track (the FX button in its header) or on the master (Audio › Master Effects). The rack is a window, so you can leave it open while you listen.

A track header showing FX 2 has two effects running. Eleven chains for voice come built in, and anything you set up can be saved beside them.

The signal path

Clip source → clip gain and fades → clip effects → summed into the track → track effects → pan and fader → summed into the master → master effects → master gain → safety limiter.

Reading the meters

Every channel has a meter down the right of its strip, and the master has the wide one across the transport. The bar is the average level; the line at its tip is the peak, held for a moment so a transient too short to move the bar still leaves a reading. The gap between them is the crest factor — on speech, normally 12 to 18 dB.

A red mark at the end means something reached full scale. Click a meter to clear it.

Export

  • Export Mixdown — the whole session as one file.
  • Export Time Selection — just the range you have selected.
  • Export Tracks Separately — one file per track, for somebody else to mix.

WAV at 16, 24 or 32-bit float, FLAC, MP3 from 64 to 320 kbps, or AAC. Sample rates from 22.05 up to 192 kHz, stereo or mono. The MP3, AAC and FLAC encoders are the ones Windows ships, so nothing extra is bundled and nothing is licensed on your behalf.

Deliver at a loudness, not near one

Set Loudness in the export dialog and Seston measures the whole mix to ITU-R BS.1770 first, then delivers it at the target with the peaks limited to the ceiling rather than clipped. −16 LUFS is podcast, −14 is what the streaming services normalise to, −19 is radio playout, −23 is EBU R128 broadcast.

Measuring means rendering twice, so it takes about twice as long. That is the whole point: loudness is a property of the whole programme and there is no honest way to know it from the first few seconds. The report at the end is measured in the file that was written — if it says −16.0 LUFS, that is what the file is.

Mono halves the file and loses nothing a single voice put there. Note that a stereo mix folded to mono is not the loudness the stereo was, so the measurement is taken after the fold, not before.

Presets

The combinations that come up are named: podcast, streaming, radio playout, broadcast, archive, voice note, master. Picking one sets the format, rate, channels and loudness together; changing any of them afterwards is allowed and moves the preset to Custom.

The ceiling only applies when the loudness is being set. Left as mixed, the export writes exactly what the mixer produced and nothing else touches it.

Settings

Edit › Preferences, or Ctrl+,.

  • Audio Hardware — output device, recording source, buffer size. The first place to look if there is no sound. MME works everywhere; WASAPI is tighter but only produces sound when the device is already running at the session's rate.
  • Auto Save — how often recovery data is written, and how many previous versions of a session file to keep.
  • Time Display — hours and minutes, decimal seconds, or samples.
  • General — snapping, follow, quick fade length, new track height, and whether to look for a newer version on start-up.

Keyboard

  • SpacePlay and pause
  • SSplit at the playhead
  • Ctrl+MMerge the selected clips
  • Ctrl+DDuplicate
  • Ctrl+TAdd a track
  • M / S / RMute, solo, arm for recording
  • G / F / EGain, fades, effects
  • LLoop the time selection
  • Shift+FFit the session on screen
  • Shift+TFit the tracks to the window
  • Ctrl+,Preferences
  • F1The full list, inside the program

When something is wrong

No sound

Ctrl+, and look at Audio Hardware. The page shows which device is actually open, which is the only way to tell whether a WASAPI request quietly fell back to MME. If in doubt, choose MME and the system default.

Nothing is being recorded

Watch the meter while recording — it shows the incoming signal, not the mix. If it is not moving, the source is wrong. On System playback, remember it captures what the machine plays, so something has to be playing.

Windows says the publisher is unknown

Seston has not been through code signing. Choose More info, then Run anyway.

Enhance Speech made it sound worse

Try Clip → Enhancement mix first and bring it down to 50 or 60% — that is immediate, and most of the time it is what a voice that already sounded reasonable wanted. If it is still wrong, run Enhance again with the speech model switched off or the strength lower; it reads the original recording, so a second pass does not add to the first. If the report said the level had to be held back, that is the sign the cleaning took out too much.