Voice SDK

Voice SDK

@sentdm/voice lets a web app place and receive Sent Voice calls in the browser, with optional React helpers. Your backend mints a short-lived voice token for each user with the Sent API. The SDK registers the browser with it, then places calls, receives them, and reports how they go. What rings when a user calls is decided by your callback, as for every call; see Answering Calls.

Installation

npm install @sentdm/voice
yarn add @sentdm/voice
pnpm add @sentdm/voice
bun add @sentdm/voice

Quick Start

Mint voice tokens on your backend

Mint tokens on your server with your Sent API key, never in the browser, for the user who is signed in:

// On your server, behind a route such as GET /api/voice-token
const response = await fetch('https://api.sent.dm/v3/channels/voice/tokens', {
  method: 'POST',
  headers: { 'x-api-key': process.env.SENT_API_KEY, 'Content-Type': 'application/json' },
  body: JSON.stringify({ identity: user.id }),
});
const { data } = await response.json();
return data.token;

identity names the user in calls: letters, digits, - and _, up to 200 characters. The user is bound to your default voice number unless you pass number, one of the numbers you turned voice on for with POST /v3/channels/voice, and the token lasts ttl seconds (600 by default, at most 3600). Return the token as text from an endpoint of your own, such as /api/voice-token. The token endpoint and the binding rules are covered in Calls From Your App.

Serve the service worker

Incoming calls reach the browser as push messages, through a service worker your app serves. Copy the one this package ships next to your pages, for example:

cp node_modules/@sentdm/voice/sw.js public/sw.js

The SDK registers sw.js relative to the page. If you serve it elsewhere, or your app already has a service worker, give it a path and scope of its own so it does not replace yours:

import SentVoice from '@sentdm/voice';

const client = new SentVoice({
  tokenProvider: () => fetch('/api/voice-token').then((response) => response.text()),
  serviceWorker: { url: '/voice/sw.js', scope: '/voice/' },
});

Register and call

import SentVoice from '@sentdm/voice';

const client = new SentVoice({
  tokenProvider: () => fetch('/api/voice-token').then((response) => response.text()),
});

// From a click handler: the browser asks for the notification permission here
await client.register();

const call = await client.connect({ to: '+14155551234' });
call.on('connected', () => console.log('Connected'));
call.on('disconnected', ({ state }) => console.log(`The call ended as ${state}`));

The first register() asks the user to allow notifications, which the browser needs to deliver incoming calls. Firefox and Safari only show that prompt after a user gesture, so call it from a click handler. It rejects when notifications are blocked or the service worker cannot be registered.

Client Options

new SentVoice(options) takes the following options. Only tokenProvider is required.

OptionDescription
tokenProviderReturns a voice token from your backend, which mints it with POST /v3/channels/voice/tokens. Called by register() and again before each token expires
registerRetriesHow many more times registering or refreshing the token is retried after a temporary failure, such as a network error. Default 2
logLeveloff, error, warn, info or debug. Default warn
loggerAn object with error, warn, info and debug methods. Default console. Voice tokens are redacted from every message
serviceWorker{ url, scope } of the worker that delivers incoming calls, a copy of @sentdm/voice/sw.js. url is resolved against the page and defaults to sw.js; scope defaults to the worker's folder
audioinputDeviceId and outputDeviceId to start with, as listed by client.audio, and element, an HTMLAudioElement that plays the other party instead of one the SDK creates
telemetry{ baseURL, disabled }. See Telemetry

Client

The client holds one registration and the calls that run on it.

PropertyDescription
stateunregistered, registering, registered, offline or destroyed. After destroy() the state is destroyed and every method throws
identityThe identity the current token was minted for. undefined until registered
numberThe number the identity is bound to. undefined until registered
callsEvery call in progress
activeCallThe call placed or answered last, until that call ends, or null
isBusyWhether activeCall is set. A pending incoming call does not make it true
audioThe microphones and speakers. See Devices
MethodDescription
register()Fetches a token from tokenProvider, registers with it, and keeps it refreshed. Resolves once the client is registered
unregister()Stops receiving calls and leaves calls in progress alone
destroy()Ends every call and tears the client down for good
connect({ to })Calls a phone number in E.164 format or another user of your app by identity. See Calls
joinConference({ name })Joins one of your account's rooms by name
EventFires
registeringregister() started
registeredThe client is registered and ready for calls
unregisteredThe client went offline: unregister() was called, or register() failed. It fires before the teardown finishes, so await unregister() before registering again
offlineThe token could not be refreshed, with a TokenExpiredError as the reason. The client keeps trying
tokenWillExpireBefore the token is refreshed, with { expiresAt }
incomingCallA call arrived, with its CallInvite. See Receiving Calls
activeCallChangedactiveCall was set or cleared, with the call or null
errorSomething failed outside a method call, such as a token refresh, with a SentVoiceError

Token Lifecycle

tokenProvider is how the SDK gets every token. It must return the token that POST /v3/channels/voice/tokens responded with, fetched fresh each time, because the SDK calls it again before each token expires.

  • register() calls it and resolves once the client is registered. The client moves from unregistered to registering to registered, emitting each state as an event.
  • A token provider that throws or rejects is retried after a short backoff, 2 more times by default (registerRetries); register() then rejects with a NetworkError. A value that is not a voice token rejects with a TokenInvalidError at once.
  • The token is refreshed at 80% of its lifetime, or 30 seconds before it expires if that is earlier, but never before half of its lifetime has passed. tokenWillExpire fires first, then tokenProvider is called with the same retries. Nothing else changes for your app.
  • If the retries run out, the client emits error with the cause and offline with a TokenExpiredError, then tries again about every 30 seconds until it is registered again; calling register() tries at once. Calls in progress go on, but connect() and joinConference() throw a NotRegisteredError while the client is offline.
  • unregister() stops receiving calls and leaves calls in progress alone. destroy() ends every call and tears the client down for good.

Tokens stay in memory and are redacted from log messages.

import SentVoice from '@sentdm/voice';

const client = new SentVoice({
  tokenProvider: async () => {
    const response = await fetch('/api/voice-token');
    if (!response.ok) throw new Error(`Fetching the voice token failed with status ${response.status}`);
    return response.text();
  },
});

client.on('registered', () => console.log('Ready for calls'));
client.on('tokenWillExpire', ({ expiresAt }) => console.log('Refreshing, expires', new Date(expiresAt)));
client.on('offline', (reason) => console.warn(`Offline (${reason.code}), trying again every 30 seconds`));

Calls

connect({ to }) calls a phone number in E.164 format, like '+14155551234', or another user of your app by identity, like 'ben'. joinConference({ name }) joins one of your account's rooms, named with letters, digits, - and _, up to 27 characters. Either throws a SentVoiceError with code INVALID_ADDRESS when to or name does not fit, and one with code CALL_IN_PROGRESS while a call is in progress or an incoming call is still waiting for an answer. Both open the microphone, and reject with a MediaPermissionError when the user refuses. to says who the user wants to reach; your backend's answer decides what rings.

A leg to a phone number, whether the answer connects the call to a number or a phone participant is added through the API, runs for at most what your account's balance affords at the destination's per-minute rate, capped at four hours. Legs to app users and rooms have no such limit because they cost nothing. A call that reaches the cap ends as completed.

const call = await client.connect({ to: 'ben' });

call.on('ringing', () => console.log('Ringing'));
call.on('connected', () => call.sendDigits('1'));
call.on('muteChanged', (isMuted) => console.log(isMuted ? 'Muted' : 'Unmuted'));
call.on('disconnected', ({ state, error }) => console.log(`The call ended as ${state}`, error?.code));

A call you place moves through initiated, ringing, answered and connected; an incoming call starts at ringing. Either kind ends as completed, failed, busy or noAnswer, which disconnected reports. client.activeCall is the call placed or answered last, until that call ends, and activeCallChanged reports each change. client.calls lists every call in progress.

PropertyDescription
idThe call id of the calling provider, for your logs. The Sent call_ id is not observable in the browser; your backend sees it on the callback question and the webhooks
directioninbound or outbound
from, to{ kind: 'user', identity }, { kind: 'number', number } or { kind: 'conference', name }
stateinitiated, ringing, answered, connected, reconnecting, completed, failed, busy or noAnswer
isMutedWhether the microphone is muted
startedAtWhen the call connected, in epoch milliseconds. undefined until then
MethodDescription
hangup()Ends the call. Does nothing once the call has ended
mute(muted)Mutes or unmutes the microphone. Toggles when called without an argument
sendDigits(digits)Sends DTMF tones down the line
getStats()Reads the call's jitter and rtt in milliseconds and its packetLoss as a fraction between 0 and 1
EventFires
ringingThe other side is being rung. Calls you place only: an incoming call starts in ringing, so it never fires this event
answeredThe other side answered
connectedAudio flows both ways
reconnecting, reconnectedThe connection dropped and came back. See Call Quality and Reconnection
muteChangedmute() changed the microphone, with the new isMuted
qualityWarningA network metric crossed its threshold, with { metric, cleared }
errorThe call failed, with a SentVoiceError. disconnected follows
disconnectedThe call ended, with { state, error }. state is completed, failed, busy or noAnswer

The SDK plays the other party through an audio element it creates. Pass your own with the audio: { element } option.

Receiving Calls

client.on('incomingCall', async (invite) => {
  console.log('Incoming call from', invite.from);
  invite.on('cancelled', () => console.log('The caller hung up'));

  // When the user answers
  const call = await invite.accept();
  call.on('disconnected', ({ state }) => console.log(`The call ended as ${state}`));

  // When the user declines instead
  // await invite.reject();
});

The microphone opens when the call arrives, and accept() asks for permission again if that was refused, then resolves with the call. When the user refuses, it rejects with a MediaPermissionError and the invite stays pending, so the user can try again. reject() declines the call, and cancelled fires when the caller hangs up first. A call that arrives while another one is in progress still raises incomingCall; your handler decides whether to accept() it or reject() it. The call accepted last becomes activeCall.

CallInvite memberDescription
from, toWho is calling and who they asked for, as addresses
statepending, accepted, rejected or cancelled
accept()Answers the call and resolves with the Call
reject()Declines the call
accepted eventThe invite was accepted, with the call
rejected eventThe invite was declined
cancelled eventThe caller hung up before an answer

Call Quality and Reconnection

qualityWarning fires when a network metric of a connected call crosses its threshold, and again with cleared: true once it recovers. Stats are sampled every half second.

metricRaised whenCleared when
jitterincoming audio jitter is over 30 ms in 3 of the last 4 samplesfewer than 3 of the last 4 are
packetLossover 1% of the incoming audio packets are lost in 3 of the last 4 samplesfewer than 3 of the last 4 are
rttthe round-trip time is over 300 msit is 300 ms or less again

The rtt warning follows every sample, so it can go on and off while the round-trip time hovers around 300 ms. A warning still raised when the call ends is not cleared.

call.on('qualityWarning', ({ metric, cleared }) => console.log(metric, cleared ? 'recovered' : 'poor'));
call.on('reconnecting', () => console.log('Connection lost, reconnecting'));
call.on('reconnected', () => console.log('Connection back'));

When the connection drops for 2 seconds, the call goes reconnecting, and reconnected fires when it comes back on the same network. A call cannot move to a new network: after a network switch its connection does not come back, and a call that has not reconnected within 5 minutes ends as failed with a NetworkError.

Handling Errors

Every error the SDK throws or emits at runtime is a SentVoiceError with a stable code, a category (auth, media, signaling, network, validation or capability) and whether it is retriable. The one exception is a programming error: using a React hook outside SentVoiceProvider throws a plain Error. Check for the subclasses with instanceof:

import { MediaPermissionError, SentVoiceError } from '@sentdm/voice/errors';

try {
  await invite.accept();
} catch (error) {
  if (error instanceof MediaPermissionError) console.warn('Allow microphone access to answer calls');
  else if (error instanceof SentVoiceError) console.error(error.code, error.message);
}
ErrorcodecategoryWhen
TokenInvalidErrorTOKEN_INVALIDauthregister(): the token provider returned something that is not a voice token
TokenExpiredErrorTOKEN_EXPIREDauththe offline event: the token could not be refreshed
NotRegisteredErrorNOT_REGISTEREDvalidationconnect() or joinConference() while not registered; registering, calling or choosing a device after destroy()
MediaPermissionErrorMEDIA_PERMISSION_DENIEDmediaconnect(), joinConference() or accept(): the microphone could not be opened
CallFailedErrorCALL_FAILEDsignalingaccept() on a call that ended before it was answered; a call that could not be set up, such as to an identity that is not registered, ends as failed with it
CallRejectedErrorCALL_REJECTEDsignalingnot raised yet: a call the other side declines ends as busy
NetworkErrorNETWORKnetworkregister(), when the token provider keeps failing or the calling engine cannot be loaded; a call whose connection failed or was lost
CapabilityUnsupportedErrorCAPABILITY_UNSUPPORTEDcapabilitysetOutputDevice() where the browser cannot choose the speaker; listing devices over plain HTTP
SentVoiceErrorINVALID_ADDRESSvalidationconnect() or joinConference() with a to or name that does not fit
SentVoiceErrorCALL_IN_PROGRESSvalidationconnect() or joinConference() while a call is in progress or an incoming call is waiting for an answer
SentVoiceErrorUNKNOWNsignalinganything else, such as register() failing because notifications are blocked

Errors that happen outside a method call arrive as events: the client's error when a token refresh fails, and a call's error, followed by disconnected with { state: 'failed', error }. The underlying engine error is not exposed on the error; the SDK reports it to Sent as telemetry. The error classes are exported from @sentdm/voice/errors and from the root entry point.

Devices

const isHeadset = (device) => device.label.includes('Headset');

const microphone = (await client.audio.inputDevices()).find(isHeadset);
if (microphone) await client.audio.setInputDevice(microphone.deviceId);

const speaker = (await client.audio.outputDevices()).find(isHeadset);
if (speaker) await client.audio.setOutputDevice(speaker.deviceId);

client.audio.on('deviceChanged', () => console.log('A device was plugged in or out'));
  • inputDevices() and outputDevices() list the microphones and speakers. Browsers leave out labels, and may list a single device, until the user has allowed microphone access in the page, which the first call asks for.
  • setInputDevice() switches the microphone of the call in progress and of later calls. While the device is missing, the default one is used.
  • setOutputDevice() plays calls through a speaker. It throws a CapabilityUnsupportedError in browsers that cannot choose the audio output.
  • Both work before register(), and take effect when it runs. To start with saved devices, pass audio: { inputDeviceId, outputDeviceId }.
  • deviceChanged fires when a device is plugged in or out.

Ringtones

The SDK plays no ringtone or ringback. Play your own on incomingCall and when a placed call is ringing, and stop it when the call is answered or ends:

const ringtone = new Audio('/sounds/ringtone.mp3');
ringtone.loop = true;

const play = () => ringtone.play().catch(() => {});
const stop = () => {
  ringtone.pause();
  ringtone.currentTime = 0;
};

// Incoming: ring until the invite is settled
client.on('incomingCall', (invite) => {
  play();
  invite.on('accepted', stop).on('rejected', stop).on('cancelled', stop);
});

// Outgoing: ringback until the call connects or ends
const call = await client.connect({ to: '+14155551234' });
call.on('ringing', play).on('connected', stop).on('disconnected', stop);

React

@sentdm/voice/react wraps one client in a provider, with hooks that re-render on its events. It supports React 18 and 19.

import { SentVoiceProvider, useActiveCall, useIncomingCall, useSentVoice } from '@sentdm/voice/react';

const getVoiceToken = () => fetch('/api/voice-token').then((response) => response.text());

export function App() {
  return (
    <SentVoiceProvider tokenProvider={getVoiceToken} autoRegister={false}>
      <Phone />
    </SentVoiceProvider>
  );
}

function Phone() {
  const { client, state, register } = useSentVoice();
  const invite = useIncomingCall();
  const { call, state: callState, isMuted, mute, hangup, duration } = useActiveCall();

  if (state !== 'registered') return <button onClick={register}>Go online</button>;
  if (invite) {
    return (
      <p>
        <button onClick={() => invite.accept()}>Answer</button>
        <button onClick={() => invite.reject()}>Decline</button>
      </p>
    );
  }
  if (call) {
    return (
      <p>
        {callState} for {duration} s <button onClick={() => mute()}>{isMuted ? 'Unmute' : 'Mute'}</button>
        <button onClick={hangup}>Hang up</button>
      </p>
    );
  }
  return <button onClick={() => client?.connect({ to: 'ben' })}>Call Ben</button>;
}
  • SentVoiceProvider takes every client option, plus autoRegister (default true), which registers as soon as the client exists. Registering needs the notification permission, which Firefox and Safari only ask for after a user gesture, so the example registers from a click.
  • The client is created after the provider mounts, so its children render first, on the server too, with client null and state 'unregistered'. Unmounting destroys the client and ends its calls.
  • Props are read once, when the client is created: remount the provider with a new key to change them, for example for another user. tokenProvider is the exception, the latest one is always used.
  • useSentVoice() gives the client, its state, register and unregister. useIncomingCall() gives the newest invite still waiting for an answer, or null. useActiveCall() gives the active call with its state, isMuted, mute, hangup, sendDigits and duration, the seconds since it connected. useAudioDevices() gives the microphones and speakers as inputs and outputs, with setInput and setOutput.
  • The hooks throw when used outside a SentVoiceProvider.
  • The entry point is marked 'use client' for the Next.js App Router.

Telemetry

The SDK reports usage and call quality data to Sent, authenticated with the voice token:

  • the SDK version, and the browser, its major version, the operating system, its version and the device type, as read from the user agent (the user agent itself is not sent)
  • how long register() took, its attempts and, when it failed, its error; how long unregister() took
  • for each call: its id, direction and outcome, how long it took to connect and how long it lasted, and at its end the average round-trip time, jitter and packet loss of samples taken every 10 seconds
  • the errors the client and its calls report: code, category, whether retriable, the message, and the underlying engine error with its name, message, stack, and properties

No audio, phone numbers or identities are added by the SDK. Batches are sent 3 seconds after the first event queued, so events that happen together share one request, and at once when a call ends, when the page is hidden and when the client is destroyed. A batch that fails, or that Sent refuses (for example with a 429 or a 5xx), is retried once, with the next batch or within 30 seconds, then dropped; telemetry never throws to your app and never delays a call. Turn it off with telemetry: { disabled: true }.

Browser Realities

  • WebRTC requires a secure context: serve your app over HTTPS, or from localhost while developing. Over plain HTTP the browser hides the microphone and service worker APIs, so registering and calling fail. It is the first thing that breaks when you open the app from another device on your network.
  • A page refresh or tab close ends any live call: WebRTC media dies with the page, and there is no resuming a call across reloads.
  • Laptop sleep or a network switch does not unregister the client: incoming calls reach it through the browser's push service, which reconnects by itself, so the client stays registered, but calls that arrive while the device sleeps or is offline are missed. If a token refresh falls in that gap, the client goes offline and tries again about every 30 seconds until it is registered again (the state goes from offline to registered).
  • One SentVoice instance per tab. The same identity registered in several tabs or devices creates one registration per device. Every device rings, the first to answer takes the call, and the others see their invite cancelled. Tabs of one browser share a registration, so they all ring.
  • Tokens are per-registration; each tab runs its own tokenProvider calls.

What the SDK Does Not Do

  • Video. Calls are audio only.
  • SIP calling. connect() takes a phone number or an identity; there is no SIP address kind.
  • Hold. There is no hold() on a call. Your callback can park a call; see Voice Recipes.
  • Transfer and participants. Adding, muting and removing participants stays on the REST API; see Calls, Recordings, and Participants.
  • Recording control. Recording is started by your callback answer or by the REST API, never from the browser.
  • Presence. The SDK does not know which identities are online. Your backend does, and your callback answer uses that knowledge.
  • Push to closed tabs. Pushes only deliver calls to open pages; the service worker never wakes a closed tab.

Source & Issues

Getting Help


On this page