Jev Ultrafast: How Speculative Multi-Head Targeting Cuts Browser Automation Latency by 25%
Hook
Three Google Flights searches in seven seconds each. No screenshots. No vision models. Just a Python agent that treats your browser's DOM like a database and makes one AI call per action instead of three.
Context
LLM-driven browser automation has a latency problem that nobody wants to talk about. The dominant pattern—screenshot the page, send the image to GPT-4V, parse out a click coordinate, execute, repeat—burns 3-8 seconds per step. Half of that is screenshot encoding and network transfer. The other half is the vision model squinting at pixels to rediscover information the browser already knows: which buttons exist, what text they contain, whether they're visible.
Jev Ultrafast exists because someone finally asked the obvious question: why are we using vision models for web pages? The DOM is already structured data. ARIA labels already encode semantics. The browser already knows what's clickable. Instead of round-tripping through image encoders, Jev reads the DOM as a structured table, sends it to a text-only model that outputs operation probabilities (CLICK, TYPE_TEXT, SELECT) plus separate target distributions for each operation type, and executes the winning combination. One network call instead of three. No image serialization. No coordinate mapping. The result is Google Flights automation that completes in seven seconds versus the typical 25-30 seconds for screenshot-based agents.
Technical Insight
The breakthrough is architectural, not algorithmic. Jev uses a speculative multi-head targeting system where a single API call to TypeSafe's Jev model returns probabilities for the operation type AND pre-computed target distributions for each operation—but only the winning operation's targets get executed. This collapses the traditional decide-what-to-do → find-where-to-do-it → generate-input pipeline into one network round trip with conditional execution.
Here's what the snapshot function in snapshot.js builds before every decision:
function snapshot() {
const elements = [];
document.querySelectorAll('a, button, input, select, textarea, [role="button"]').forEach((el, idx) => {
if (el.offsetParent !== null) { // visible check
elements.push({
id: idx,
tag: el.tagName,
text: el.innerText?.slice(0, 100) || '',
aria: el.getAttribute('aria-label') || '',
type: el.type || '',
value: el.value || '',
href: el.href || '',
_ref: el // keep DOM reference for execution
});
}
});
return elements;
}
This runs atomically in the browser and returns a flat array of interactive elements with their semantic metadata. No repeated getProperty calls. No stale selector resolution. The _ref field keeps a live JavaScript reference to the actual DOM node so execution can happen without re-querying.
The TypeSafe API receives this structured state plus the task goal and returns a response like this (simplified):
{
"operation_probs": {
"CLICK": 0.73,
"TYPE_TEXT": 0.18,
"SELECT": 0.06,
"DONE": 0.03
},
"click_target": [0.02, 0.89, 0.05, 0.01, ...], # distribution over element IDs
"type_text_target": [0.91, 0.03, 0.02, ...],
"select_target": [0.15, 0.78, 0.04, ...],
"reasoning": "Need to click the departure city field"
}
Notice the three separate target heads. The model computes distributions for click targets, type targets, and select targets simultaneously, but agent.py only executes the head matching the winning operation. If CLICK wins (0.73 probability), the executor samples from click_target, ignoring the other two distributions entirely. This is speculative execution at the decision level—you're computing multiple futures but only committing to one.
Text generation is deferred to a second, smaller model and aggressively cached:
if op == "TYPE_TEXT":
target_element = elements[sampled_target_id]
cache_key = (goal, target_element['aria'], target_element['text'])
if cache_key in text_cache:
text_to_type = text_cache[cache_key]
else:
text_to_type = mercury_model.generate(
f"Field: {target_element['aria']}. Goal: {goal}. Type:"
)
text_cache[cache_key] = text_to_type
This matters during navigation retries. If the user clicks a link and the page reloads, the TYPE_TEXT decision gets re-evaluated on the new DOM. But if the input field still exists with the same ARIA label and goal, the cached text gets reused without re-generating. This survives stale-page errors that would otherwise force redundant LLM calls.
Execution validation happens at the last possible moment in browser.py. The agent doesn't check if an element is clickable during planning—it checks during execution:
def click(self, element_ref):
# Freshness check
if element_ref.ownerDocument != self.page.document:
raise StaleElementError("Page changed since planning")
# Geometry check
rect = element_ref.getBoundingClientRect()
if rect.width == 0 or rect.height == 0:
raise NotClickableError("Element has zero size")
# Occlusion check
center_x, center_y = rect.left + rect.width/2, rect.top + rect.height/2
topmost = self.page.document.elementFromPoint(center_x, center_y)
if topmost != element_ref and not element_ref.contains(topmost):
raise OccludedError(f"Element covered by {topmost.tagName}")
element_ref.click()
self._adaptive_wait(element_ref.tagName)
The occlusion check uses elementFromPoint to find what's actually under the click coordinate. If it's not the target element or a child of the target, another element is blocking the click—probably a modal or loading spinner. This catches the common failure mode where an element exists in the DOM but isn't actually interactive.
Adaptive waiting is tied to interaction type, not blanket sleeps:
def _adaptive_wait(self, interaction_type):
if interaction_type == 'SELECT': # combobox suggestions
time.sleep(0.2) # 200ms for AJAX
elif interaction_type == 'CLICK':
time.sleep(0.05) # 50ms for CSS animations
# TYPE_TEXT gets zero wait
These timeouts are hardcoded based on observed browser behavior, not learned. The logs timestamp execution before waiting, so reported decision latency doesn't include DOM settling time. This makes benchmarks honest—you're measuring the agent's think time, not how long comboboxes take to populate.
The entire agent loop in agent.py is under 300 lines because there's no vision pipeline, no screenshot encoding, no selector generation, and no retry logic for stale elements (they just raise and the outer loop re-snapshots). The DOM is the ground truth. The model consumes structured data. Execution is validated once, at commit time.
Gotcha
The MVP DOM reader only handles standard HTML and ARIA controls. Shadow DOM is invisible to querySelectorAll, so custom web components (most modern design systems) won't appear in the snapshot. Iframes require separate snapshots per frame, which isn't implemented. Canvas-based UIs and file upload dialogs are completely unsupported because they're not representable as DOM elements with text/ARIA metadata. If your target site uses Material-UI, Salesforce Lightning, or any framework that renders custom controls via shadow roots, Jev will see an empty page.
The hardcoded 200ms/50ms wait times assume typical server response and animation speeds. Sites with slower backends or unusual loading patterns will either waste time (if you wait too long) or miss state changes (if you don't wait long enough). There's no dynamic adjustment based on observed latency. The single-browser-profile architecture means cookies and localStorage bleed across tasks—if you automate login on Site A, then run a task on Site B, Site B sees Site A's session state. You need manual profile management for isolation, and there's no built-in cleanup between runs.
Most critically, there's no outcome validation. The DONE operation fires when the agent thinks it's finished, but there's no automatic verification that the task actually succeeded. The Google Flights example in flights.py manually checks that the route, dates, and results list match expectations after the agent reports completion. You're writing task-specific assertions anyway, which means Jev is a step executor, not a goal achiever.
Verdict
Use if: You're automating standard web forms or search interfaces where latency matters—credential stuffing for security research, OSINT workflows on government databases, high-volume data extraction from structured sites. The 25% speed improvement over screenshot-based agents is real, and the sub-10-second execution opens workflows that weren't viable at 30-second latencies. Also use if you're researching action-space compression for embodied agents and want to study how speculative multi-head prediction reduces round trips. Skip if: Your target sites use shadow DOM, modern JavaScript frameworks with custom components, or require multi-tab coordination. Skip if you need production reliability—three successful runs isn't a benchmark, and the lack of built-in verification means you're writing test harnesses anyway. Skip if your automation needs to handle file uploads, canvas interactions, or anything beyond clicking/typing into standard HTML controls. For those cases, use screenshot-based Browser Use (35% slower but handles complex UIs) or raw Playwright with GPT-4o for field extraction (100x faster once mapped but requires per-site setup).