Starlog — Page 66
// LATEST
AI Agents
Building Security-Aware AI Assistants with VirusTotal and the Model Context Protocol
Cybersecurity
CVE-Bench: Testing Whether AI Agents Can Actually Hack
Developer Tools
ELT-Bench: The First Realistic Benchmark for AI Agents Building Data Pipelines
AI Agents
Codex CLI: OpenAI's Rust-Powered Terminal Agent That Brings ChatGPT to Your Command Line
AI Agents
Roo Code: A Multi-Modal AI Agent Architecture for VS Code
Developer Tools
Hyperscan: How Intel Matches 50,000 Regex Patterns at 10+ Gigabits Per Second
AI Agents
VulnBot: When Multi-Agent LLMs Take Over Penetration Testing
AI Agents
smolagents: Why Hugging Face Built an Agent Framework in Just 1,000 Lines
Cybersecurity
CAI: The Uncensored AI Framework Rewriting the Rules of Offensive Security
AI Agents
OpenManus-RL: Teaching LLM Agents to Think Better Through Reinforcement Learning
AI Agents
DeerFlow: ByteDance's Production-Grade Framework for Hour-Long Autonomous AI Agents
AI Dev Tools
Inside Microsoft's AI Red Teaming Playground: Training Security Professionals to Break LLMs
AI Agents
Robin: Building a Multi-Agent System That Generates Drug Discovery Hypotheses
AI Agents
WebVoyager: Teaching GPT-4V to Navigate the Web Like a Human
Cybersecurity
HackBench: Measuring What Happens When LLMs Learn to Exploit Vulnerabilities
AI Agents
Inside the Daily Knowledge Engine Tracking 2,000+ Autonomous Agent Papers
AI Agents
Autono: Why Dynamic ReAct Beats Static Planning for Failure-Prone Agent Tasks
Data & Knowledge
Persona-Hub: How Tencent's Billion-Scale Perspective Engine Reimagines Synthetic Data
AI Agents
ToolHive: Bringing Kubernetes-Grade Security to Model Context Protocol Servers
LLM Engineering
AgentDojo: The Security Benchmark That Exposes LLM Agents' Achilles Heel
AI Agents
Ruflo: Building Self-Learning Agent Swarms for Claude with Federation
AI Dev Tools
Strudel: How TidalCycles' Pattern Algebra Was Reimagined for the Web
AI Agents
Magentic-UI: Microsoft's Plan-Then-Execute Web Agent That Shows Its Work
AI Agents