> your AI agent picks dependencies from memory; give it dated facts — try starlog.dev ↗ vet your agent's deps ↗ vibe-coding is fine. vibe-importing isn’t. — try starlog.dev ↗ vibe-importing isn’t fine ↗ your agent has never seen your private packages — try starlog.dev ↗ facts for private packages ↗ a linter for the dependencies your AI agent picks — try starlog.dev ↗ a linter for agent deps ↗ whois is redacted, cdns mask the rest — get the real operator — try whoisgeni.us ↗ who really runs that domain ↗ domain attribution that shows its work — full evidence chain — try whoisgeni.us ↗ domain intel w/ evidence ↗

Starlog — Page 69

// LATEST

Developer Tools

BountyBench: A Framework for Benchmarking AI Agents on Security Vulnerability Research

★ 88 Jupyter Notebook May 8, 2026
AI Agents

Inside the Framework Measuring How Good AI Agents Are at Hacking

★ 8 Python May 8, 2026
AI Agents

WASP: The Security Benchmark That Catches What Your Web Agent Misses

★ 87 Python May 8, 2026
AI Agents

Web-Shepherd: Training Web Agents with Process Rewards Instead of Binary Success

★ 56 Python May 8, 2026
AI Agents

GUARDIAN: Detecting When Your AI Agents Start Lying to Each Other

★ 8 Python May 8, 2026
AI Agents

Process-Supervised RL for Agentic RAG: How ReasonRAG Achieves 18x Data Efficiency

★ 14 Python May 8, 2026
AI Agents

Building a Unified AI Gateway: How IBM's ContextForge Federates MCP, REST, and Agent Protocols

★ 3.8k Python May 8, 2026
AI Agents

AgentAuditor: The Invisible Research Project That Might Transform AI Agent Verification

★ 4 May 8, 2026
Cybersecurity

RAPTOR: Building an Autonomous Security Agent from Claude Code and Adversarial Thinking

★ 2.8k Python May 8, 2026
AI Dev Tools

Happy: Monitoring AI Coding Agents From Your Phone Without Leaking Your Code

★ 21.4k TypeScript May 8, 2026
LLM Engineering

TheAgentCompany: The First Real-World Benchmark That Makes AI Agents Look Bad

★ 715 Python May 8, 2026
Developer Tools

Training Web Agents Through Test-Time Interaction: Inside TTI's Filtered BC Approach

★ 75 Python May 8, 2026
Developer Tools

AGI SDK: Building Browser Agents Against Production-Quality Web Replicas

★ 409 Python May 8, 2026
AI Agents

SPORT: Teaching Multimodal Agents to Self-Improve Without Human Labels

★ 20 Python May 8, 2026
AI Agents

RF-Agent: Teaching Language Models to Design Reward Functions Through Tree Search

★ 11 Jupyter Notebook May 8, 2026
LLM Engineering

SEC-bench: A NeurIPS Framework for Benchmarking LLM Agents Against Real Security Vulnerabilities

★ 77 Python May 8, 2026
AI Agents

Superpowers: Teaching AI Agents to Stop Coding Like Caffeinated Interns

★ 214.0k Shell May 8, 2026
Automation

Stagehand: The Browser Automation SDK That Caches AI Actions Like Code

★ 22.9k TypeScript May 8, 2026
Automation

Steel Browser: The Open-Source Browser API That Lets AI Agents See the Web

★ 7.1k TypeScript May 8, 2026
Cybersecurity

HackingBuddyGPT: Teaching LLMs to Think Like Penetration Testers

★ 1.1k Python May 8, 2026
AI Agents

ARTEMIS: Stanford's Multi-Agent Red Teaming System That Orchestrates LLMs to Hunt Vulnerabilities

★ 516 Rust May 8, 2026
AI Agents

LatentMAS: How Multi-Agent Systems Learned to Think Without Speaking

★ 966 Python May 8, 2026
AI Dev Tools

HumanLayer: The Context Engineering Framework That's Mostly Vapor

★ 10.9k TypeScript May 8, 2026
AI Agents

Maestro: Orchestrating Multiple AI Coding Agents with Git Worktrees and Batch Automation

★ 3.0k TypeScript May 8, 2026