speechout — n_pbt_speechout #
← Component reference · Guide contents
Reading aloud: the application reads a text sentence by sentence and tells you where it has got to. No third-party service, no API key — the synthesis is the workstation's own.
▶ See it live — Demo application, Speech out tile: the preview, the code behind it and this page, side by side.
At a glance #
| Nonvisual object | n_pbt_speechout |
| Used for | Making a text heard: accessibility, busy hands, a notification nobody is looking at |
| Principle | You hand over the text; the component cuts it into sentences and tells you which one it is reading |
| Dependency | The workstation's speech synthesis — no third-party service, no API key |
Quick start #
// once, when the window opens
inv_voice.of_open()
// then, wherever you need it
inv_voice.is_lang = inv_voice.LANG_FR_FR
inv_voice.of_speak("Bonjour. Votre commande est expediee.")
Nonvisual: wiring the events #
A voice has nothing to show. So the component draws nothing: no player bar to find room for, no space taken from what it serves.
And you have nothing to wire for it: a nonvisual object has no window, hence no doorbell, but the component pulls its own events on the PowerBuilder loop while the voice is open and raises them on the object. You only write the ue_* handlers (no receiver, no timer).
// Just speak -- the ue_* events arrive on their own :
inv_voice.of_speak("Good morning. Your order has shipped.")
// the ue_sentence / ue_word / ue_stopped events arrive on their own
Following the reading in YOUR text #
ue_sentence carries the index and the text of the sentence being read. That is where the highlight comes from — in your mle_, your datawindow or your statictext: you know where your text is, we never did.
The way back exists too: of_speak_from() resumes at one precise sentence, which is what you wire to a click on a paragraph. The numbers come from ue_sentence, so they always point at what was actually read.
The cutting is the component's, not yours:
of_sentence_count()returns its own count. Do not recount on your side — the two would drift.
What the workstation can really say #
of_languages() returns the languages this workstation can actually pronounce, without duplicates. That is the question a user asks: not "which voices exist", but "is my language there".
of_voices() goes one step down and names the voices themselves. Both lists come from the machine, not from us: never hard-code a name.
Voice recognition has no equivalent, and that is not an oversight: recognition carries no list of the languages it accepts.
Without any component of yours, gnv_utils.of_speech_languages(as_tags[]) gives the same list — the DLL asks a hidden voice, then drops it: that is what a dialog asks before the voice exists, which language to offer, which one to read in. The call is synchronous and can take up to two seconds the first time: the voice list arrives late, and the call waits for it. And gnv_utils.of_speech_voices(as_names[], as_langs[]) gives the voices themselves with their language, to offer "Hortense" or "Julie" rather than a tag; gnv_utils.of_locale_name(as_tag) gives a tag its readable name — "French (France)" for fr-FR.
string ls_tags[]
if inv_voice.of_languages(ls_tags) > 0 then inv_voice.is_lang = ls_tags[1]
Properties #
| Property | Type | Default | Role |
|---|---|---|---|
is_lang | string | en-US | Language read aloud, in BCP-47 (LANG_* constants). Decides which voice is picked — while is_voice is empty, a voice name pinning its own language. With no voice for that language the workstation reads with the one it has, and ue_voice_fallback names both |
is_text | string | "" | The text to read. The component cuts it into sentences; three tags say HOW to read a piece: [pause=500], [spell]…[/spell], [say-as=digits]…[/say-as] (see below) |
is_voice | string | "" | Voice name, taken from of_voices(). Empty = the first one that speaks is_lang |
ii_rate | integer | 100 | Rate, in PERCENT of the normal rate (10 to 400). Changed mid-reading, it applies to the next sentence |
ii_pitch | integer | 100 | Voice pitch, in PERCENT of the normal pitch (0 to 200) |
ii_volume | integer | 100 | Loudness in percent (0 silent to 100), the video player's scale |
il_timeout_ms | long | 300000 | Longest a SYNCHRONOUS reading may last: five minutes. Past it, of_speak_sync returns -4 and the reading is cut (is_last_error says why) |
is_last_error | string | "" | Why the last of_speak_sync returned -4: voice not created, engine failure, timeout |
Methods #
| Method | Role |
|---|---|
of_open ( ) | Creates the voice. Optional — of_speak does it — but calling it when the window opens pays the cost once, out of the way of the first sentence. Returns a positive number once the voice exists, 0 or less when it could not be created |
of_is_open ( ) | True once the voice exists |
of_speak ( string as_text ) | Sets the text and reads it, from the first sentence. Without an argument, re-reads is_text. Returns 0 once sent, -1 when the voice could not be created |
of_speak_from ( long al_index ) | Resumes the reading at one precise sentence. Returns 0 once sent, -1 when the voice could not be created |
of_pause ( ) | Suspends the reading where it is. Returns 0 once sent, -1 when the voice could not be created |
of_resume ( ) | Picks up where of_pause left off. Returns 0 once sent, -1 when the voice could not be created |
of_stop ( ) | Stops the reading; of_speak starts again from the first sentence |
of_is_speaking ( ) | True while a sentence is being read — a paused reading still counts. Asked to the component, never a stale copy |
of_sentence_count ( ) | Returns how many sentences the component made of the text |
of_count ( ) → integer | How many sentences the component made of the text — the same number of_sentence_count answers. The library asks this question under one name everywhere |
of_voices ( ref string as_names[] ) | Fills the array with the voices installed on this workstation and returns how many |
of_languages ( ref string as_tags[] ) | Fills the array with the languages this workstation can pronounce, without duplicates, and returns how many |
of_voice_used ( ) | The voice the component will actually hand to the engine — not always the one is_lang asked for: a workstation carries the voices someone installed on it and no others. Empty = the component imposes none, and the engine takes its own, the voice of the system language. Naming that one would be a guess. ue_error says the same when reading starts; this reads it beforehand |
of_speak_sync ( { string as_text } ) | Reads and WAITS for the end: the next line runs after the last sentence, the window keeps painting. Returns 0 once ended, -4 on failure or timeout (is_last_error) |
of_enqueue ( string as_text ) | QUEUES a text: read at once when the voice is idle, after the current reading otherwise, never cutting it. Returns 0, -1 when the voice could not be created |
of_clear_queue ( ) | Forgets the waiting texts, without cutting the one being read. Returns 0, -1 when the voice could not be created |
of_queue_count ( ) | Returns the number of texts still waiting (the one being read is not counted) |
of_add_replacement ( string as_from, string as_to ) | A pronunciation rule: every WHOLE word as_from is read as as_to (PB → PowerBuilder). Applied before the tags, kept by the object. Returns 0, -5 when as_from is empty |
of_clear_replacements ( ) | Empties the dictionary. Returns 0 |
of_replacement_count ( ) | Returns the number of rules in the dictionary |
of_pick_voice ( string as_lang, string as_gender ) | Picks an INSTALLED voice for a language and, when one exists, a gender (GENDER_FEMALE, GENDER_MALE, GENDER_ANY): exact language, then its family. Puts it in is_voice and returns it; empty when no voice speaks that language |
of_duration ( ) | Returns the ESTIMATED number of milliseconds the reading will take (words per minute at the requested rate, pauses included): for a progress bar, not a stopwatch. The pace LEARNS the voice: every sentence read to its end measures the real one, kept per voice on this workstation |
of_position ( ) | Returns the estimated number of milliseconds already read, refined by the words the engine reports; 0 when nothing is being read |
of_progress ( ) | Returns the estimated progress, 0 to 100 |
of_spoken_text ( ) | The sentences AS THE VOICE GETS THEM, one per line (dictionary applied, tags resolved): the text to display to follow word by word |
of_process_events ( ) | Drains the queued events and raises them on this object. The component's own pump calls it for you while the voice is open — you never call it |
of_close ( ) | Releases the voice, stopping whatever it was saying first. The destructor calls it |
of_reset ( ) | Puts every property back to its original value |
Events #
| Event | Fired when |
|---|---|
ue_started (string as_lang) | The reading begins; as_lang recalls in which language |
ue_stopped ( ) | The last sentence is done, or of_stop was called |
ue_paused ( ) | The reading is suspended |
ue_resumed ( ) | The reading picks up again |
ue_sentence (long al_index, string as_text) | For each sentence, with its rank and its text: that is how you follow the reading elsewhere in the window |
ue_error (string as_message) | The workstation has no engine or no voice at all, or the voice fails. A voice merely MISSING is not an error: that is ue_voice_fallback |
ue_voices_ready (long al_count) | The engine has filled its voice list — it arrives late; of_voices, of_languages, of_voice_used and of_pick_voice wait for it by themselves (2.5 s at most), so this event only says WHEN it came; al_count says how many the workstation has |
ue_word (long al_index, long al_start, long al_length) | The WORD being said in sentence al_index: Mid(sentence, al_start, al_length). When the engine reports words (most Windows voices) |
ue_queue_done ( ) | The last text of the queue (of_enqueue) has been read |
ue_voice_fallback (string as_wanted, string as_used) | The voice or language asked for is not on this workstation; as_used names the one reading instead. Information, not an error: the reading goes on |
Pronunciation: pauses, spelling, dictionary #
WebView2 has no SSML. So the text carries three tags of its own, resolved before the cut into sentences, and a dictionary of whole words applied before them. An unknown tag is read as written.
| Tag | Effect |
|---|---|
[pause=500] | A 500 ms silence (10 s at most). The pause ends the current sentence |
[spell]ABC12[/spell] | Every character said one by one: "A, B, C, 1, 2" |
[say-as=digits]4152[/say-as] | The digits one by one, not "four thousand one hundred fifty-two" |
[say-as=characters]…[/say-as] | Same as [spell] |
// The dictionary : whole words, in the order added, case-sensitive
inv_voice.of_add_replacement("PB", "PowerBuilder")
inv_voice.of_add_replacement("Mme", "Madame")
inv_voice.of_add_replacement("4152", "[say-as=digits]4152[/say-as]") // a rule may add a tag
inv_voice.of_speak("Mme Durand, PB order 4152 [pause=600] code [spell]PBT[/spell].")
Queue, synchronous reading, progress #
of_speakcuts,of_enqueuewaits. An application that announces events (an alert, a result, a notification) queues: two announcements close together are both heard, andue_queue_donesays when the last one is read.of_clear_queueforgets what waits without cutting;of_stopdoes both.of_speak_syncreturns at the end. The script waits for the last sentence (the window keeps painting), bounded byil_timeout_ms, which CUTS the reading when it is reached — for "read this, then ask the question".of_duration,of_position,of_progressare estimates: the engine says nothing about length, they count words at the requested rate, refined by the words the engine reports. Enough for a progress bar read from a timer, not for a stopwatch.of_pick_voice(language, gender)picks an installed voice by language then gender, and puts it inis_voice. The gender comes from the first name Windows gives each voice; a voice with an unknown first name answers toGENDER_ANY.- No audio export. WebView2's synthesis renders no stream: the reading cannot be written to a file. It is not a missing setting, it is the engine.
Examples #
Reading a notification #
inv_voice.of_speak("Order 4152 has been shipped. It arrives on Thursday.")
Lighting up the current sentence #
// in ue_sentence, on the nonvisual object
st_read.text = as_text
// and a click on a paragraph resumes from it
inv_voice.of_speak_from(2)
Choosing the voice and the rate #
string ls_voices[]
if inv_voice.of_voices(ls_voices) > 0 then inv_voice.is_voice = ls_voices[1]
inv_voice.ii_rate = 75
inv_voice.of_speak()
Announcing without cutting: the queue #
// Each event of the application is queued : all of them are heard, in order
inv_voice.of_enqueue("Order 4152 has been shipped.")
inv_voice.of_enqueue("Order 4153 is ready.")
// ue_queue_done fires after the last one
Lighting up the current word #
// in ue_sentence : keep the sentence
is_sentence = as_text
// in ue_word : the word is Mid(is_sentence, al_start, al_length)
st_read.text = Left(is_sentence, al_start - 1) + "[" + Mid(is_sentence, al_start, al_length) + "]" + Mid(is_sentence, al_start + al_length)
A voice by language and gender, then reading while waiting for the end #
inv_voice.of_pick_voice(n_pbt_speechout.LANG_FR_FR, n_pbt_speechout.GENDER_FEMALE)
if inv_voice.of_speak_sync("Please confirm the order.") = 0 then
li_answer = MessageBox("Order", "Confirm ?", Question!, YesNo!)
end if
Good practice #
- One sentence at a time: the component hands the engine a single sentence and chains. That is what allows a clean stop between two sentences, and what avoids Chromium quietly giving up past about fifteen seconds.
- Light up the sentence being read in your text, from
ue_sentence— that is half of what reading aloud is for. - Call
of_languages()rather than assuming: a language installed on your machine may not be on the customer's. - Do not call
of_speakin a loop over a result set: each call cuts the previous one, and the user hears only beginnings. That is whatof_enqueueis for.