Skip to content

Build1 publisher2 min readPublished

GitHub Security Lab credits step-by-step LLM taskflows with 24 Android vulnerability reports

GitHub Security Lab says its open-source LLM taskflows have found and reported 24 Android app vulnerabilities by auditing code in staged steps. Teams with a GitHub Copilot license can point the same audit at their own repositories.

The Engineer · Build desk

Illustration accompanying GitHub Security Lab credits step-by-step LLM taskflows with 24 Android vulnerability reports

What happened

  • One Android stage hands the model a list of vulnerability classes to check per entry point, such as confused deputy and insecure broadcasts for intent handlers.
  • The OsmAnd navigation app, with over 10 million Android downloads, yielded three of the findings, one of which lets malicious apps track the device's location.
  • Setup is a codespace started from the seclab-taskflows repository and one command, run_mobile.sh, pointed at the target repo.
  • The post's author reported more than 20 of the vulnerabilities personally, and GitHub lists new disclosures on its advisories page.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

  • cost The strict prompt guards against missed bugs only when it runs repeatedly, and each extra run draws more premium requests and tokens, so how deep the audit goes depends on the team's budget.
  • decision The has_vulnerability checkmarks are the model's verdicts, so a team adopting the tool has to assign someone to confirm each one before anything is filed as a report.
  • capability The vulnerability checklist lives in an open-source YAML taskflow, so a team can add the bug classes its own apps tend to have and rerun the audit against them.

The entry-point split is the part I would copy first. A taskflow called gather_mobile_entry_point_info.yaml takes the entry points and separates them into mobile and non-mobile [10]. A repository can hold an Android app next to web servers and desktop clients, and the model still works against the correct attack surface for each [10]. The Android stages build on general audit taskflows that the author's colleagues Peter and Mo wrote earlier [9].

The checklist stage exists because of a constraint the author states outright. Mobile vulnerability classes are less widely known, and LLMs are non-deterministic, so the taskflow makes sure the model checks certain essential classes [12]. I think a checklist suits Android, where the common bug classes can be written down in advance. A written list does not depend on the model remembering them on any particular run.

Coverage then comes from repetition. The post runs the strict prompt alongside a broader one across multiple runs. The author wrote that "the strict prompt and repeated runs ensure obvious vulnerabilities aren't missed, while the broad prompt lets the AI apply its creativity to the fullest." [13] Every one of those runs is billed. The prompts use premium model requests, and GitHub warns that the many tool calls can easily consume a large amount of tokens [5][6]. One pass over a medium-sized repository might take an hour or two [8].

The location-tracking bug in OsmAnd starts at the kind of surface the checklist is built for. The app exports an activity called MapActivity, which handles opening settings files and deeplinks [15]. Components outside the app can launch an exported activity [15]. When it opens a settings file, the app accepts intent extras including settings_version, silent_import and replace [16].

Whether the 24 carries over to someone else's code depends on the target. The checklist only helps if an app exposes the entry point types it names. The team also has to run the audit enough times to smooth out the model's variance [11][12]. GitHub's figure is a count of vulnerabilities its own researchers found and reported [3]. The post does not include a single-prompt baseline or a false-positive count. That leaves the claim that staging finds bugs the model "would have missed entirely" resting on GitHub's own account [2].

What to watch

  • New entries on GitHub Security Lab's advisories page credited to the taskflows, moving the count past 24.
  • A published comparison of taskflow runs against a single-prompt audit on the same Android repositories, with false-positive counts.
  • Per-run premium request and token figures from teams running run_mobile.sh on their own code.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories