Files
knowledge-wiki/channels/1518733120831226028.md
T

2.9 KiB

Channel Wiki: #jarvis-jr-test-001

Channel ID: 1518733120831226028 Last updated: 2026-07-31 13:27 UTC

Purpose

Voice-enabled test channel for conversations with Hermes/Jarvis.

Voice attribution and responsiveness

  • Multiple people can join and speak in this voice session.
  • Incoming transcribed speech may be labeled as RootAtSkic even when another participant is speaking.
  • On 2026-07-28, Linas and Martynas joined and spoke to Jarvis, but Jarvis incorrectly treated their turns as Lego/RootAtSkic.
  • Do not assume the session-level user ID identifies the speaker of every voice-transcribed turn.
  • When speaker identity materially matters and is not explicit, infer from the conversation carefully or ask who is speaking.
  • In a live voice conversation, acknowledge a request immediately before loading skills or calling tools.
  • If processing may take more than a brief moment, tell the person or group what is being checked and that it may take a little longer.
  • Give short progress updates during extended work rather than leaving participants in silence.
  • CONFIRMED voice-response requirement (Lego, 2026-07-31): generate human-friendly responses that are clear and easy to follow. Expect the human participant to guide the conversation and ask follow-up questions; answer conversationally rather than delivering dense written-style output.
  • Treat long runs of identical phrases such as repeated “I'm sorry” as probable STT/transcription loops; acknowledge the likely glitch without treating every repetition as intentional.
  • Root cause confirmed on 2026-07-28: the local Whisper-compatible STT endpoint intermittently emitted 614- and 472-character repeated “I'm sorry” decoder loops from Discord voice audio. Silence/noise retests returned empty transcripts, so the failure is intermittent.
  • Profile-local plugin /opt/data/plugins/voice-improvements/ is enabled. It adds generic repeated-phrase rejection (four or more repetitions covering at least 80% of words) and progressive Discord TTS.
  • Progressive TTS splits long responses at sentence/word boundaries into at most four chunks (target 380 characters), synthesizes them sequentially in a worker, and plays each chunk as soon as ready. Real Piper verification produced 3 non-empty chunks; first ready in 1.4 seconds, all complete in 4.5 seconds.
  • The plugin activates after the Hermes gateway/session is restarted.
  • Activity-based presence is configured for voice channel 1518735773657206835 (jarvis-jr-test-001) linked to text channel 1518733120831226028. Jarvis auto-joins when a human enters or is already present at gateway startup, stays while humans remain, and auto-leaves 30 seconds after the last human departs. A rejoin during the grace period cancels the leave. Core /voice join and /voice leave remain available as manual controls.

People observed in this voice chat

  • Lego / RootAtSkic
  • Linas (linas_02251, Discord ID 1483820991720329227)
  • Martynas