> your AI agent picks dependencies from memory; give it dated facts — try starlog.dev ↗ vet your agent's deps ↗ vibe-coding is fine. vibe-importing isn’t. — try starlog.dev ↗ vibe-importing isn’t fine ↗ your agent has never seen your private packages — try starlog.dev ↗ facts for private packages ↗ a linter for the dependencies your AI agent picks — try starlog.dev ↗ a linter for agent deps ↗ whois is redacted, cdns mask the rest — get the real operator — try whoisgeni.us ↗ who really runs that domain ↗ domain attribution that shows its work — full evidence chain — try whoisgeni.us ↗ domain intel w/ evidence ↗

← Back to Articles

SCOUT-2: Building a Desktop AI Assistant with Asynchronous Cognitive Operations

[ View on GitHub ]

SCOUT-2: Building a Desktop AI Assistant with Asynchronous Cognitive Operations

Hook

Most AI assistants forget who you are the moment you close the window. SCOUT-2 runs a background 'cognitive operations' service that continues thinking about your conversations after they end, quietly building a persistent model of your context.

Context

The proliferation of LLM providers has created a new problem: vendor lock-in at the API level. Build your assistant on OpenAI's function calling, and migrating to Anthropic means rewriting your tool integration layer. Build on Anthropic's prompt caching, and you can't A/B test against Google's Gemini without maintaining parallel codebases.

SCOUT-2 tackles this by implementing a provider abstraction layer that treats OpenAI, Anthropic, Mistral, Google, and HuggingFace as interchangeable backends. But it goes further than simple API normalization—it introduces 'cognitive operations,' a background service that asynchronously processes conversations to extract insights, generate titles, and update user profiles. This creates a primitive form of long-term memory where the assistant actually learns from your interaction history, rather than treating each conversation as an isolated context window. The medical persona support, including EMR storage and NCBI API integration, suggests this emerged from healthcare research needs where provider flexibility and persistent patient context are critical.

Technical Insight

Backend Services

Managers

user input

persist/retrieve

load definitions

map to

switch between

audio pipeline

spawn task

generate titles

update profiles

extract info

inject context

response

render

audio playback

orchestrates

tkinter GUI Thread

main.py Async Context

ConversationManager

ProviderManager

ToolManager

PersonaManager

UserDataManager

Cognitive Operations Service

LLM Providers

OpenAI/Anthropic/Mistral

Google Cloud TTS/STT

SQLite DB

chat history

JSON Tool Definitions

function_maps

Python Callables

System architecture — auto-generated

The architecture centers on five manager classes that handle orthogonal concerns. The ProviderManager abstracts LLM, TTS, and STT services, while ToolManager loads JSON function definitions and maps them to Python callables. What makes SCOUT-2 interesting is how these managers coordinate through async context managers in main.py:

@asynccontextmanager
async def lifespan(app):
    # Initialize all managers
    conversation_mgr = ConversationManager(db_path)
    provider_mgr = ProviderManager()
    tool_mgr = ToolManager(function_maps_dir)
    
    # Start cognitive operations background service
    cognitive_task = asyncio.create_task(
        cognitive_operations_service(
            conversation_mgr,
            provider_mgr
        )
    )
    
    yield {
        'conversation': conversation_mgr,
        'provider': provider_mgr,
        'tool': tool_mgr
    }
    
    # Cleanup
    cognitive_task.cancel()
    await conversation_mgr.close()

The cognitive operations service is where things get clever. After each conversation completes, it spawns async tasks to generate conversation titles using the LLM and update user profiles based on extracted information:

async def cognitive_operations_service(conv_mgr, provider_mgr):
    while True:
        # Check for conversations without titles
        untitled = await conv_mgr.get_untitled_conversations()
        
        for conv_id in untitled:
            messages = await conv_mgr.get_messages(conv_id)
            
            # Use LLM to generate a title from conversation context
            title_prompt = f"Generate a 3-5 word title for this conversation: {messages[:3]}"
            title = await provider_mgr.complete(title_prompt)
            
            await conv_mgr.update_title(conv_id, title)
        
        # Extract user profile updates from recent conversations
        await update_user_profiles(conv_mgr, provider_mgr)
        
        await asyncio.sleep(30)  # Run every 30 seconds

This background processing means SCOUT-2 is always working even when idle—conversations get automatically titled, user preferences get extracted and stored, and the system builds a richer context model over time.

Function calling is implemented as a translation layer rather than using provider-native APIs. The ToolManager loads JSON schemas from the tools/ directory and converts them into provider-specific formats at runtime. For example, an OpenAI function definition gets translated to Anthropic's tool format when you switch providers:

def translate_tool_schema(tool_json, target_provider):
    if target_provider == 'anthropic':
        return {
            'name': tool_json['name'],
            'description': tool_json['description'],
            'input_schema': {
                'type': 'object',
                'properties': tool_json['parameters']['properties'],
                'required': tool_json['parameters'].get('required', [])
            }
        }
    elif target_provider == 'openai':
        return tool_json  # Already in OpenAI format
    # ... other providers

The persona system is implemented as dynamic prompt injection. The PersonaManager maintains JSON files defining different assistant personalities, and the UserDataManager injects user-specific context (profile data, medical records, system info) into each conversation:

def build_system_prompt(persona_name, user_data):
    persona = load_persona(persona_name)  # Load base persona JSON
    
    system_prompt = persona['base_prompt']
    
    # Inject user context
    if user_data.get('medical_records'):
        system_prompt += f"\n\nPatient Context: {user_data['medical_records']}"
    
    if user_data.get('preferences'):
        system_prompt += f"\n\nUser Preferences: {user_data['preferences']}"
    
    return system_prompt

The medical persona is particularly sophisticated—it integrates with NCBI's E-utilities API to fetch research articles and stores patient data in SQLite as EMRs. This treats user context as a domain-specific language that gets compiled into prompts, allowing the assistant to maintain medical context across sessions.

The tkinter GUI runs synchronously while the backend operations are async, bridged through asyncio.run() calls wrapped in threading. Voice input/output uses Google Cloud's STT/TTS APIs, but audio playback happens synchronously in the GUI thread, which can freeze the interface during longer audio clips.

Gotcha

The Windows-only deployment with admin privilege requirements immediately limits the audience. The batch script launcher expects Visual C++ build tools and specific Windows APIs, making this a non-starter for Linux or Mac users. Cross-platform support would require rewriting the audio pipeline and removing Windows-specific dependencies.

The custom function calling implementation, while providing consistent behavior across providers, creates ongoing maintenance debt. As OpenAI, Anthropic, and Google evolve their native function calling APIs with better streaming support, caching, and parallel tool execution, SCOUT-2's translation layer will fall behind. You're trading provider-agnostic interfaces today for performance limitations tomorrow.

The medical features are a regulatory minefield. Storing patient EMRs in unencrypted SQLite without HIPAA compliance documentation, audit logging, or data retention policies makes this dangerous for actual healthcare use. The NCBI integration is read-only research retrieval, which is safer, but the EMR storage suggests clinical use cases that this codebase isn't legally prepared for. The custom license prohibiting redistribution means you can't fork it to add compliance features without negotiating terms with the author.

Conversation history uses SQLite without vector search or semantic retrieval. As your conversation database grows, finding relevant past context becomes a linear scan problem. There's no embedding generation, no similarity search, no RAG pipeline—just chronological append-only storage.

Verdict

Use if: You're a Windows power user experimenting with multi-provider LLM interfaces and want to see how background cognitive services can create persistent memory without blocking the main interaction loop. The medical persona features are valuable if you're doing healthcare research (not clinical deployment) and need NCBI integration. The provider abstraction shows how to build vendor-agnostic AI tools when you need runtime provider switching. Skip if: You need cross-platform support, production-ready medical compliance, or plan to fork/redistribute the code (the custom license blocks this). The lack of vector search makes this unsuitable for long-term conversational AI where semantic retrieval matters. If you want native function calling performance or a modern web UI instead of tkinter, look at Open-WebUI or LangChain-based alternatives instead.