The approach
Two open-source Python packages that handle that layer once: visvoai-ai turns one id string into a streaming model across Gemini, Anthropic and OpenAI-compatible APIs, with a live model registry and cost metering, and visvoai-core is a lean LangGraph agent loop that applications extend through hooks instead of forking. A terminal coding agent and a multi-user reference platform both run on them unmodified.
How it works
- Model layer:
visvoai-aiturns one id string into a streaming model for any supported provider, with a live registry of prices, context windows and thinking levels. - Agent runtime:
visvoai-coreruns the agent and tools loop on LangGraph, with a soft step cap and hooks for extending it. - Tool retrieval: for large tool sets, only the tools a step needs are retrieved and bound (BM25, or hybrid with embeddings).
- Applications: a terminal coding agent and a multi-user reference platform, both built on the packages without forking them.
Open-source packages
- visvoai-ai: pass one id like gemini:gemini-2.5-flash and get a streaming LangChain model, across Gemini, Anthropic and any OpenAI-compatible endpoint (OpenAI, Together, Groq, OpenRouter built in).
- A live model registry (price, context window, capabilities, thinking levels): 44 built-in models, extendable to about 4,000 from models.dev with an offline fallback, plus cost_of and usage_from for metering and one 4-level thinking scale over 8 provider mechanisms.
- visvoai-core: a LangGraph agent loop with a soft step cap that forces a clean final answer, 8 runtime hooks to extend nodes, routing, state and checkpointing without forking, and lifecycle persistence hooks.
- One tool contract that mixes plain typed functions, class-based tools and LangChain tools in a list, plus per-round BM25 + embedding tool retrieval for large tool sets.
- CI enforces the boundary between the public packages and the private platform; 104 + 42 tests, MIT, Python 3.11+.
Install and first call, from the READMEs:
pip install "visvoai-ai[gemini]" visvoai-core
from visvoai.ai import build_chat_model
model = build_chat_model("gemini:gemini-2.5-flash")
for chunk in model.stream("Explain attention in one sentence."):
print(chunk.content, end="")
from visvoai.ai import build_chat_model
from visvoai.core import AgentRuntime, ask
graph = AgentRuntime().build_graph(
model=build_chat_model("gemini:gemini-2.5-flash"),
core_tools=[list_dir, read_file], # plain typed functions with docstrings
system_prompt="You help explore a codebase.",
)
print(await ask(graph, "What does this directory contain?"))
Terminal coding agent
- A terminal coding agent (Textual UI) built on the two packages, unmodified, and the reference for how far they go; 607 tests.
- Shell commands are sorted into read or write: reads run without prompts inside a no-write OS sandbox (macOS sandbox-exec, Linux bubblewrap), and changes go through an approval gate with three modes and path confinement.
- Subagents run in parallel with isolated context, live logs and per-run cost traces, alongside on-demand skills and MCP servers; anything a cloned repo defines stays off until you approve it.
- Each tool batch and turn is snapshotted to a shadow git repo, so /rewind restores the code, the conversation or both to before any earlier question.
pip install visvoai-cli
export GEMINI_API_KEY=...
visvoai
Reference platform
A private, multi-user reference platform (agent harness) built on the same packages to test and benchmark them at scale; it is not publicly hosted.
- Dynamic tool binding so agents stop loading every tool, which took ~57% of a typical 46k-token prompt: a small core is always bound and the rest are retrieved per step, with Gemini-embedded MCP tools.
- Rewrote tool descriptions to lift recall@8 from 91% to 97% on a 211-case benchmark over 35 built-in tools; a later run over 38 built-in and 73 MCP tools reached 100% recall@8 with hybrid retrieval.
- Agent harness and loop on LangGraph: parallel tool batches, repeat-call blocking, schema-guided recovery from bad arguments, step caps and context compaction.
- ~40 built-in tools with role-scoped availability and resource access checks, plus human-in-the-loop approvals (approve, edit, approve-all, reject) that pause and resume the agent.
- User-defined agents, 17 on-demand skills, and subagents with isolated context, depth caps and parallel runs that stream live to the parent's UI.
- An MCP client and OAuth connectors written from scratch (Google, Microsoft, Slack, Notion, HubSpot) that sync server tools by content hash, repair malformed arguments and move large results to files.
- RAG over a per-user document library: PDF/DOCX/PPTX parsing, heading-aware chunking, Gemini embeddings and hybrid BM25 + vector search in Weaviate with citations, plus long-term memory.
- Versioned artifacts that agents and users co-edit with conflict checks, SQL over spreadsheets in DuckDB, and a network-isolated sandbox for running code.
- Per-user data isolation and encrypted per-user API keys, LLM cost and token tracking with a trace viewer, automatic audit logs, conversation branching, and stream resume over Redis Streams.
