PORTFOLIO PROJECT · CASE STUDY

From Article to Autonomous Digital Presenter: Building an End-to-End AI Video Platform at Near-Zero Infrastructure Cost

I designed and built this platform — architecture, pipeline, governance model and the two Digital Presenters who front it.

ALEXANDROS and LYDIA are AI-powered Digital Presenters created for THE ARCHON. They do not represent real people or employees.

ALEXANDROS & LYDIA in conversation

Why this exists

THE ARCHON publishes executive technology analysis — articles, frameworks, guides, infographics — through Portfolio and Technology Intelligence. Written analysis has a ceiling: a reader has to arrive, already know they want to read, and commit the time. Video reaches a different audience, on a different platform, with a different attention economy — but producing it the conventional way (a presenter, a studio, an editor, repeated for every article) does not scale against a one-person publishing operation.

The brief I set myself was narrow and specific: take an existing published article — or, separately, a manually authored script with no source article at all — and produce a presenter-led video of it, without a camera, a studio, or a recurring per-video cost, while keeping a human decision in the loop before anything reaches a public audience.

ALEXANDROS & LYDIA

ALEXANDROS and LYDIA are THE ARCHON's two AI-powered Digital Presenters. Each has a fixed visual identity, a locked voice contract and a governed pronunciation layer, so the same name is never pronounced two different ways across two videos. They are disclosed as AI-powered digital presenters everywhere they appear — on their public profile pages, in video descriptions, and in this case study — never presented as real people.

Most content is SOLO: one presenter, one topic, derived from one article. A second format — DUO Conversations — puts both presenters in the same scene, addressing a topic from two perspectives (pictured above). DUO is deliberately a conversation formatbetween the two existing presenters, never a third identity: there is no “third presenter,” only ALEXANDROS and LYDIA, together.

The end-to-end pipeline

Every video — SOLO or DUO, article-derived or manually authored — moves through the same pipeline. A DUO Conversation does not require a source article at all; Title, Summary, Technology Domain(s) and Tags can be entered directly.

Article or Manual Script→Presenter Selection→Kokoro TTS→Pronunciation Layer→MuseTalk Render→SOLO or DUO Orchestration→Private Storage→Human Review→Approval→YouTube (Unlisted)→Portfolio / TI Discovery

Script, voice and pronunciation governance

Speech is synthesized with Kokoro, an open-weights text-to-speech model — no per-character metered API, no vendor lock-in on the voice itself. Each presenter has one locked voice contract (ALEXANDROS: am_adam; LYDIA: af_heart), so voice identity never drifts between videos.

Pronunciation of proper nouns is a real production problem with an open-weights TTS model, and I treat it as a governed input, not a one-off fix: a single, tested, word-boundary-safe override function rewrites specific tokens before synthesis — for example an explicit IPA transcription for “ALEXANDROS” — applied identically whether the presenter is speaking SOLO or as one half of a DUO Conversation. The override changes only what Kokoro hears; the canonical script text a viewer would read is never altered.

Presenter render: MuseTalk and DUO orchestration

The talking-avatar video itself is generated with MuseTalk 1.5, a real-time lip-sync model, driven from a small set of owner-approved master portraits per presenter. SOLO rendering animates one presenter against their own master image. DUO rendering is a harder problem: both presenters occupy one frame, and the pipeline has to route each spoken segment to the correct half of the scene — ALEXANDROS consistently on the left, LYDIA consistently on the right — while keeping the crop geometry of each presenter's face region tuned independently, since a crop setting that looks right on one side of the frame can clip the other presenter's mouth.

That tuning was genuinely iterative: real GPU renders, visual review, and small, evidence-based parameter adjustments rather than guesses — the kind of work that does not show up in an architecture diagram but is most of what makes a DUO video look like two presenters in conversation instead of two presenters pasted into the same frame.

Private storage, human review and approval — a feature, not a gap

Every rendered video lands in private object storage first. Nothing is uploaded to YouTube automatically. A human review and approval step sits between rendering and publication for every single video, with no bypass path — I did not build this platform to remove human judgment from what gets published under THE ARCHON's name; I built it to remove the manual production labor before that judgment point, so the review is the only step that still requires me.

This is a deliberate governance choice, not a limitation I plan to engineer away. A platform that published automatically the moment rendering finished would be faster and would also be the wrong platform to have built.

Publication and discovery

Approved videos publish to YouTube as unlisted — discoverable through THE ARCHON's own surfaces, not indexed as public search content in their own right. Portfolio and Technology Intelligence each surface the two presenters and DUO Conversations on their own “Meet the Digital Presenters” sections and on a dedicated ALEXANDROS & LYDIA Conversation Hub, and eligible Conversations participate in the same Technology Intelligence corpus, Business Needs classification and Ask Dimitrios AI knowledge layer every other content format already uses — never a second, parallel discovery system built just for video.

Near-zero infrastructure cost

The platform is deliberately built around an open-stack, cost-conscious architecture: an open-weights TTS model instead of a metered voice API; MuseTalk instead of a commercial avatar-rendering service; rendering dispatched to a cloud GPU notebook runtime instead of a dedicated, always-on GPU server; and private object storage billed by usage rather than reserved capacity. None of this is mathematically free — compute, storage and hosting all have a real, if small, marginal cost per video — but the architecture was chosen specifically so that producing one more video is a fraction of what a managed, camera-and-studio equivalent would cost, without giving up render quality or editorial control.

Operational lessons

  • Pronunciation is a governance problem, not a one-time fix — it needed its own locked, tested override layer.
  • DUO framing (two presenters, one scene) is a genuinely harder rendering problem than two independent SOLO videos, because a crop or timing adjustment that helps one presenter can visibly hurt the other in the same frame.
  • The human approval gate is not friction to engineer out — it is the single point where editorial judgment stays in a pipeline that automates everything before it.
  • Reusing one shared eligibility query, one shared discovery corpus and one shared digest processor for every content format — including Conversations — kept the system from growing a second, parallel implementation for every new format.
  • Automated production artifacts (like a generated YouTube thumbnail) need the same review discipline as the video itself — an early DUO thumbnail was generated from two individual portraits side by side instead of a real shared scene, and needed a targeted fix so future Conversations render correctly without manual correction.

Current status

The platform is live. ALEXANDROS and LYDIA each have published SOLO briefings. The official “Meet ALEXANDROS & LYDIA” video is the presenters' introduction — not itself a thematic Conversation, and excluded from the Conversation count by design. The first genuine thematic DUO Conversation, “Why Is Technology Still a Male-Dominated Profession?”, has been generated, reviewed, approved and published. This page reports only what is currently verified — specific production-volume, timing or audience metrics will be added once there is enough real data to report, not before.

← Back to Digital Projects