Product & Engineering Portfolio

A look at the products and systems built at AccelerMedia, including App9, ChatBench, Jaxia, Singagram.ai, CallOnline, and the software that keeps them running.

An app you wished existed.A song only you could give.

Create an app by describing it. Send a birthday song full of inside jokes. Give your customers an assistant that knows your business.

I’m Jacob, AccelerMedia’s Founder & CTO. I design these experiences and build the AI, integrations, and services that make them possible.

App9 Builder interface for describing and creating an application

App9 Builder

Live

That app idea has somewhere to go.

Describe a booking tool, customer portal, or dashboard. Build it through conversation, try it in a live preview, and publish it with code you can export to GitHub. Generated apps can connect to App9’s accounts, AI credits, and customer records.

Generated code gets its own sandbox. React and TypeScript apps run in isolated preview containers. App-specific credentials control access to shared services. Automated browser tests exercise generated apps, capture failures, and check whether a follow-up edit breaks an earlier feature.

More engineering detail

Generation runs as a phased plan. The agent drafts an application blueprint, generates it in stages, runs it, and reads the results before deciding what to change. Loop and completion detection stop it from re-editing the same file or declaring the app finished early, and a separate debugging pass reads runtime errors, screenshots, and console output before proposing a fix.

Every preview is a container with its own lifetime and credentials. Preview access uses revocable tokens validated independently of the management API, and the port a preview exposes is bound to a signed token, so a guessed URL reaches nothing. Prompt-injection defenses treat generated content and third-party text as data. Rate limits are enforced on two separate backends so a burst of requests cannot exhaust a single store.

Generated apps act on behalf of their owner. Access grants decide which shared services an app may call, the owner’s credit balance pays for its AI usage, and hosted form submissions become tagged CRM records for the owner. The builder exposes the same operations over MCP with OAuth so other agents can drive it, and each project’s code lives in a Git repository the owner controls.

Quality is measured. A benchmark harness builds a set of reference apps, opens each one in a real browser, exercises its workflow, and records source snapshots, screenshots, errors, and timing. The same runner replays earlier scenarios after a follow-up edit to catch regressions, and static analysis plus an automated review run on every pull request.

Explore App9 Builder
Singagram characters and personalized song greeting creator

Singagram.ai

Live

A gift that sings their story.

Put their name, your favorite memory, and that ridiculous inside joke into a birthday song or affectionate roast. Choose the musical feel, review the lyrics, and add a singing character. The recipient gets a greeting they can replay and share.

Music, animation, and checkout in the same creative flow. The system coordinates lyric drafts, song generation, character video, and delivery. Customers can revisit their choices while retaining approved lyrics and job progress. Web and native clients share TypeScript contracts for the creation steps and job states.

More engineering detail

Each greeting is a resumable job with explicit stages: lyric drafting, song generation, character video, finishing, and delivery. Credit checks happen before every paid stage, and a job that fails or resumes keeps the outputs of completed stages so nothing already paid for is repeated. Web and native clients share the same TypeScript contracts for creation steps and job states.

Video rendering runs on a GPU fleet with a control plane. Workers send heartbeats and claim jobs matched to their capabilities; leases expire when a worker disappears, attempt counters bound retries, and four priority tiers keep interactive customer work ahead of batch rendering. The fleet spans private GPU hardware for baseline capacity and cloud-hosted serverless GPUs for overflow, and failover between them is a deliberate operator decision, which keeps a local outage from turning into a surprise bill.

Timing is deterministic. A performance-window algorithm aligns the generated video to the first reliable sung lyric, plans holds, crossfades, and secondary footage around it, and a single finalizer produces the canonical output so two renders of the same job agree.

The singing performance is generated as video from the selected character artwork and finished song. The renderer uses the character image as a reference and the song audio to drive the performance; the finishing pipeline aligns the result to lyric timing, adds captions and branding, and produces the shareable video. Supported custom-character generations can use a customer-uploaded reference image through the same video-generation path. Synthetic production monitors place a real order on a schedule and alert on any stage that stops completing.

Make a Singagram
App9 Post staging demo workspace with scheduled posts, approvals, inbox, and delivery results
Current staging workspace · demonstration data

App9 Post

Live

Get the post out. Keep the final say.

Prepare a campaign once, adapt the copy for each account, and review it before it goes out. Scheduling can find the next suitable slot for each channel independently, so one busy account doesn’t hold up the others.

Every destination gets its own delivery history. The scheduler, API, TypeScript SDK, and MCP tools share approval rules. Queue-based delivery tracks retries and partial success per account, so an interrupted attempt can be reconciled without blindly publishing the campaign again.

More engineering detail

Delivery is queue-based and per destination. Each account has its own schedule, receipts, and retry state, so a partial failure is reconciled against what actually went out. Durable Objects coordinate provider rate limits per account while separate queues handle publishing, media, and webhooks.

Agents get the same controls as people. Any agent-initiated action goes through a prepare, approve, execute sequence: the prepared action is fingerprinted, a person approves it, and a short-lived one-time token authorizes exactly that action once. The scheduler, API, TypeScript SDK, and MCP tools all enforce the same approval rules.

A media service handles the work a serverless runtime cannot. It detects musical downbeats, scores candidate cut points by harmonic and loudness continuity, optionally compares video frames across the seam, and renders exact loops with FFmpeg. Its jobs use scope-restricted leases so host-side work and container work never claim each other’s tasks.

Adaptive campaigns can rank approved clips using verified performance evidence; that capability is gated separately and rolls out per account. Publishing, inbox, and analytics coverage are verified per provider before they are offered.

Explore App9 Post
Jaxia customer support interface with a shared inbox and human handoff

Jaxia

Live

Pick up the conversation, right where they left off.

A visitor can search your help center, ask the assistant, and bring in your team within the same conversation. Lead details and customer history stay attached when a person takes over.

The help article and the AI answer start with the same knowledge. Jaxia’s RAG workflow retrieves approved, published content with audience permissions applied. Jaxia coordinates tool calls and records execution traces for review. Configurable AI budgets reserve usage before generation and hand the conversation to a person when the allowance is exhausted.

More engineering detail

The assistant answers from the same approved knowledge as the help center. Articles move through draft, publish, index, and ready states; publishing writes to a durable outbox that indexes the content asynchronously, and a retrieval allowlist decides whether an article may reach customer-facing or internal-only conversations.

Retrieved text is treated as data. Web results, site grounding, and support articles are fenced inside the prompt with explicit instructions that nothing inside the fence can change the assistant’s rules, and uploaded documents are fenced the same way. The contract test suite asserts these fences on every change.

Spend is reserved before generation. Each scope has an AI budget; a reservation is taken when the conversation starts a model call and settled afterward, and when the allowance is exhausted the conversation hands to a person with the transcript, lead details, and customer history attached. Tool calls and model calls record execution traces with timing and token usage for review.

Jaxia also runs a shared work board. People and agents claim the same work items; agent claims carry capability tags, leases, and heartbeats, delegation depth is limited, and outcomes hand back through the same records people use. The web widget and the Expo client reuse one set of support APIs and one conversation history.

Meet Jaxia
ChatBench interface comparing answers from different AI models

ChatBench

Live

Give the same problem to a different mind.

Compare several AI answers to your own question, with response time, token use, and credit cost alongside them. Blind Arena hides model names so you can judge the answer before seeing who wrote it.

Comparisons built around both quality and cost. The Next.js application streams provider responses and records their usage. Blind comparisons shuffle the model order and keep it stable across reloads. Model availability is checked against a launch gate that requires both a response and successful usage settlement.

More engineering detail

Every comparison is a measured run. The application streams responses from several providers at once, records tokens, latency, and credit cost per response, and keeps text already shown to the user if a stream fails midway. A launch gate requires a model to both respond and settle usage before it is offered.

Blind Arena shuffles the model order with a stable seed so a reload cannot reveal who wrote what. Shared comparisons retain the result metadata needed to inspect a run later.

Benchmarks drive routing. An intake process scores models across dimensions, an automated judge grades answers for accuracy, reasoning, and clarity, and a routing engine compiles those scores into a versioned table. The router classifies each request deterministically, without calling a model, and selects the cheapest option that clears the quality floor for that request’s dimensions, returning a fallback chain and the savings against the frontier model. The shared AI gateway consumes the published table for its automatic model aliases.

A leaderboard publishes the current scores, and a soak harness exercises the realtime path under sustained load before releases.

Compare models in ChatBench
CallOnline voice agent dashboard showing calls and usage

CallOnline

Live

Let your agent make the call.

Give an AI workflow a phone connection, then bring back the transcript, summary, and structured outcome. For incoming calls, Jaxia can answer as an AI receptionist; missed-call text-back gives the caller another way to reach the business.

A conversation has to handle interruptions. Bidirectional WebSocket audio connects live calls to speech recognition and generated voice responses. The system stops playback when a caller interrupts and maintains context throughout the conversation. I’ve tested the Jaxia/CallOnline receptionist with real inbound calls in production.

More engineering detail

Audio is bidirectional and streamed. Live calls connect over WebSocket to streaming speech recognition and generated voice; playback stops the moment a caller interrupts, and the agent keeps its context across the interruption. Recording consent, do-not-call lists, quiet hours, and two-party-consent jurisdictions are enforced by a compliance policy before the agent speaks.

Call minutes are metered like AI tokens. A credit reservation is taken when a call starts and committed or rolled back against the carrier’s final call status, and every usage event is deduplicated so a repeated webhook cannot bill twice. Missed calls trigger a text-back with the same duplicate protection.

Agent actions carry approvals. Transfers, bookings, and outbound follow-ups go through the same prepare, approve, execute controls as the rest of the platform, and each call produces a transcript, summary, and structured outcome that flows into support and CRM records. REST and MCP expose calling to other applications and agents. Voice output is tiered across providers so a business chooses the balance of cost and naturalness.

Explore CallOnline

The platform underneath these products.

Eight products share one account system, one AI gateway, one memory service, and one way of running expensive compute. The decisions below are what let a small team operate all of them.

One gateway for every model call

Every product reaches AI through one hosted gateway, so routing, spending, and credentials are decided once.

More engineering detail

The gateway speaks an OpenAI-compatible interface and routes across several commercial providers and self-hosted models. Allowlists decide which models each product may use, aliases resolve to concrete models, and a fallback chain retries on rate limits and provider errors.

Spend is reserved before work starts. A request reserves credits, the call runs, and the reservation is committed or rolled back with an idempotency key so a retried request cannot charge twice. The same protocol governs LLM tokens, agent workflow steps, and telephony minutes, all settling against one account ledger, and cached input tokens are metered separately so cost attribution stays accurate.

Products never hold provider keys. Credentials and OAuth tokens are stored encrypted in the account service and brokered to the gateway per request, and an owner grants permission before local or bring-your-own-key execution is allowed. Threshold-triggered auto-reload keeps a workspace funded within bounds the owner sets.

Memory and retrieval that respect tenants

Agents remember across sessions and search a workspace’s knowledge without ever seeing another workspace’s data.

More engineering detail

Knowledge is stored as versioned Markdown and split into sections; each section gets one vector embedding, and a lexical index sits beside the vector index. A search runs both, fuses the ranked lists with reciprocal rank fusion. When embeddings are unavailable, search degrades to lexical-only and keeps working.

Tenant isolation is enforced twice: the vector query filters on workspace metadata, and every database read joins on the workspace, so a cross-tenant hit is impossible even if one layer were misconfigured. Re-indexing is incremental; only sections whose text changed are re-embedded, which keeps embedding cost proportional to edits.

Retrieval quality is measured with an evaluation harness that scores reciprocal rank, hit rate, and NDCG across lexical, vector, and fused strategies. Long-term memory is curated: daily notes and candidate facts are promoted into durable memory by a weighted score, and an active-memory pack assembles what an agent needs within a token budget.

Inference where the workload belongs

Privacy-sensitive work runs on private GPU hardware, bursts run on cloud-hosted serverless GPUs, and the gateway routes between them.

More engineering detail

Private GPU systems on-premises serve open-weight language models plus video and speech models for workloads where data must stay in-house. Cloud-hosted serverless GPUs absorb overflow. The gateway’s local broker dispatches to private hardware with per-host authentication and a credit hold per dispatch, and failover to the cloud tier is a deliberate operator decision, so a local outage becomes a visible choice instead of a silent bill.

Batch rendering runs on a control plane shared with the products: capability-matched claims, leases, heartbeats, attempt counters, and priority tiers keep interactive work ahead of background jobs across both tiers.

The economics are explicit. Suitable workloads on private hardware avoid an estimated several thousand dollars a month in equivalent hosted API usage, and every model call, workflow step, and render carries its cost back to the account that requested it.

Identity, billing, and the operational floor

One sign-in, one credit balance, and the same operational controls under every product.

More engineering detail

The account service is an OAuth 2.0 authorization server with social sign-in, WebAuthn passkeys, session management, and published authorization-server metadata, so products and third-party integrations authenticate the same way. Teams and entitlements are multi-tenant, and a credit ledger unifies app-store and card purchases into one balance.

The service tier is serverless: dozens of edge services with durable state, queues, object storage, vector search, workflows, and containers, plus Docker-based GPU workers where a serverless runtime cannot do the job. Rate limiting runs on two backends, webhooks are verified with signatures, and every mutating integration endpoint keeps a request-scoped replay ledger so a retried call is idempotent.

Operations are the same for every product: separate staging and production deploys, automated secret rotation, fleet-wide error monitoring, synthetic production monitors that exercise real customer flows on a schedule, and a library of written operating procedures the agent workforce follows.

Other problems I’m working on.

  • Site9 brings WordPress changes into a preview, approval, and rollback flow, alongside SEO and article-writing operations. Live.
  • Sell9 connects WooCommerce orders, marketplace inventory, and print-on-demand fulfillment. Its order workflow handles repeated events without deducting stock or starting fulfillment twice. Live.
  • CallAnywhere puts international calling quotes and spending limits in a browser dialer, with provider selection by cost and connection quality. Live.
  • Luxium is a semi-modular software synthesizer written in C++: a compiled per-sample voice graph, six analog-modeled filters, expressive per-note control and microtuning, measured anti-aliasing and performance gates, and signed offline licensing with a transactional outbox.

Looking for someone who can lead the team and get into the code?

These projects show how I approach product design, engineering, and day-to-day operations. If your team is hiring, or you have a product you want to build, I’d love to hear about it.

Talk with Jacob

Growth has an operating cost. I design for that, too. The platform section above shows the mechanisms: spending reserved before work starts, jobs that resume without repeating paid stages, approvals that make agent actions idempotent, and expensive compute placed by workload.

I also operate 100+ WordPress sites reaching 200,000+ monthly viewers. AccelerMedia’s game portfolio has passed 500,000 installs. See the media and games work.