BuildIndependently confirmed2 publishers3 min readPublished
Hugging Face's ML-intern agent trains small custom models only after the user approves a budget
Hugging Face's ML-intern agent turned HuggingChat prompts into published models, among them a $1.90 citrus-disease vision model and a $16 prompt rewriter. Spend stays at zero until the user approves a cap, so a small model becomes a priced request.
The Engineer · Build desk

What happened
- On 335 test photos, the fine-tuned Qwen3.5-2B citrus model named the right problem 52.8% of the time after two epochs on one A10G, up from the base model's 14.9%.
- The 0.8B Pocket Rewriter, distilled from Qwen-Image 2.1's 9B rewriter that needs about 20 GB of memory, runs on a CPU and returns valid output 99.7% of the time on a quarter of the tokens.
- Among the other builds are two Qwen-Image 2.1 LoRAs, one for camera angles and one for swapping out doodles, plus a 4-step distillation of a 260M-parameter text-to-image model.
- Compute for all six projects came to about $103, according to the dev.to account of the post.
- The author's prompts require a smoke test before the full run, such as 50 training steps followed by a check that the saved weights actually changed.
Why it matters
- cost A budget planned from the two widely quoted runs will undershoot; the six-run record supports planning nearer $20 per model, before any data preparation is counted.
- decision Because no paid job starts without explicit approval, a custom model becomes a spend request that whoever holds the cap can approve or refuse before GPU time is billed.
- constraint Teams without someone who tracks trainer releases and open issues will be writing the thinner, riskier brief, so the agent narrows the staffing need for small models without removing it.
- exposure In the author's runs every model ended up public on the Hub with its evaluation, so a team training on private data has to settle where outputs land before it approves the budget.
The cap is a sentence in the prompt. One of the author's reads: "Cap total spend at USD 12 and ask me before exceeding it." [3] A line of text cannot stop a GPU job by itself. According to the post, the limit holds because the agent starts each task with a zero-dollar budget and needs permission before it executes any paid job [2]. Leave the cap out and it proposes a couple of paths sized to the project, then asks which one the user wants [4].
The two prices both write-ups repeat sit at the cheap end. Subtract them from the six-project total reported by dev.to and the other four runs come to $85.10, or about $21 each [23]. At least one of those four cost more than the $16 rewriter [23]. The $1.90 is a compute figure [5]. Its 3,017-image dataset was assembled beforehand by the author, with Claude, from three Project-AgML sources [6]. The rewriter's $16 covers making its own data: inside that budget, the 9B teacher labelled 8,797 example requests [8].
What stayed expensive is the brief. "The first message is where I spend my effort," wrote the post's author, a Hugging Face team member [11][18]. Briefs name the dataset, the base model and the training script, and put anything already checked under a heading that says "Verified facts, do not re-derive" [20]. For the camera-angle LoRA, that section recorded which trainer had just added transparent-image support and which open GitHub issues made the fallback trainer risky [21]. Knowing either is an ML engineer's job. The intern wants a long ticket: the first prompt ran about 450 words and the sixth about 2,000 [10].
A thinner brief still worked. The citrus prompt had no verified-facts section, and the fine-tune still beat its Qwen3.5-2B base by about 3.5 times on accuracy [14][24]. That comparison exists because the prompt asked for it: "Also report the base model's zero-shot score on the same metric before training so we can see the gain." [12] Against its own baseline the gain is real. The same model is wrong on about 47% of its test photos [25].
A human still has to decide when a model is good enough, and the agent's evaluation output helps. On a Huggy LoRA trained on FLUX.2 klein base 4B from 84 captioned drawings, it saved a checkpoint every 100 steps and drew the same prompts with each [22]. Step 200 was the first fully on-model, and from step 500 the style began bleeding into other prompts [22].
In our view the budget gate is the part other agent builders should copy, because a reviewer sees a price before any hardware is billed [2]. The brief that keeps the price small still takes someone who knows the training stack. ML-intern is live in HuggingChat behind a mode switch [17], and the author's seven prompts for the six models are on GitHub at yvrjsharma/ml-intern-prompts [19].
What to watch
- Independent reruns of the published prompts that reproduce, or miss, the 52.8% citrus score and its $1.90 compute bill.
- Per-run cost and failure records from users other than the author, showing whether roughly $20 a model holds with shorter briefs.
- Whether ML-intern lets a person other than the prompt writer approve spend, as a team purchasing process would require.
Clarity's read
What the record supports and how the coverage leans. The claims behind it follow.
Reality
- Evidence50
- Adoption
- Insufficient
- Hype gap+20
- Incentives80
- Confidence60
Claim ledger
Ranked by verification strength, evidence, and original report placement.
- [1]
ML-intern is an agent mode inside HuggingChat that takes a prompt describing a model, plans the project, asks for a spending limit before running anything, trains and evaluates the result against a baseline, and publishes it on the Hub with its evaluation in the model card.
- [2]
ML-intern begins every task with a zero-dollar budget and needs permission before executing paid jobs; the post says this is why the spending limit stays strictly enforced.
- [3]
"Cap total spend at USD 12 and ask me before exceeding it."
ReportedSupportedSource: Prompt line written by the post's author2 sources— create a free account to open themView cited source - [4]
When a prompt leaves out a budget, the agent suggests a couple of paths depending on project size and asks which the user prefers.
- [5]
On 335 test photos, the base Qwen3.5-2B model named the right citrus problem 14.9% of the time; after two epochs on one A10G the fine-tuned model got 52.8%; compute cost was about USD 1.90.
- [6]
The author used Claude to put together a training dataset merged from three Project-AgML sources on the Hub: 3,017 annotated images across 21 pests, illnesses, nutritional gaps and treatment approaches; ML-intern fine-tuned Qwen3.5-2B on it.
- [7]
The 0.8B Pocket Rewriter, a small version of Qwen-Image 2.1's 9B prompt rewriter that needs about 20 GB of memory, runs on a CPU, returns valid output 99.7% of the time and uses about a quarter of the teacher's tokens.
- [8]
The rewriter project's compute, including having the 9B model label 8,797 example requests, came to USD 16.
- [9]
Total compute across all six projects came to about $103.
ReportedSupportedSource: dev.to account of the Hugging Face post2 sources— create a free account to open themView cited source - [10]
The author's first prompt was about 450 words; by the sixth project it was closer to 2,000.
- [11]
"The first message is where I spend my effort."
ReportedSupportedSource: The post's author, a Hugging Face team member2 sources— create a free account to open themView cited source - [12]
"Also report the base model's zero-shot score on the same metric before training so we can see the gain."
ReportedSupportedSource: Line from the author's citrus prompt2 sources— create a free account to open themView cited source - [13]
For the image LoRAs the author asked for 50 training steps and then a check that the saved weights had actually changed, before paying for the full run.
- [14]
The citrus brief had no verified-facts section, and ML-intern still produced a model that more than tripled the accuracy of the Qwen3.5-2B model.
- [15]
Other examples include camera-angle and doodle-replacement LoRAs for Qwen-Image 2.1 and a 4-step distilled version of a 260M-parameter text-to-image model.
- [16]
ML-intern trains, evaluates and publishes on Hugging Face hardware, and each of the author's projects ended as a public model on the Hub with its evaluation in the model card.
- [17]
ML-intern is available now inside HuggingChat by switching on the mode and submitting a prompt.
- [18]
The post was written by a Hugging Face team member and covers six models built over a few days.
- [19]
All seven of the author's prompts are on GitHub at yvrjsharma/ml-intern-prompts, exactly as written.
- [20]
The author's prompts name the dataset, the base model and the training script, and put anything already checked under a heading that says "Verified facts, do not re-derive".
- [21]
For the camera-angle LoRA, the verified-facts section listed which trainer had just added transparent-image support and which open GitHub issues made the fallback trainer risky.
- [22]
For a Huggy LoRA on FLUX.2 klein base 4B trained on 84 captioned drawings, the agent saved a checkpoint every 100 steps and drew the same prompts with each; step 200 was the first fully on-model, and from step 500 Huggy's style started bleeding into other prompts.
- [23]
The four projects other than the citrus model and the rewriter cost about $85.10 combined, about $21 each on average, so at least one cost more than $16.
- [24]
The fine-tuned citrus model's accuracy is about 3.5 times the base model's.
- [25]
The fine-tuned citrus model names the wrong problem on about 47% of the 335 test photos.
Sources
2 independent publishers whose own reporting we read for this story.
- dev.toHuggingChat's ML-intern builds custom models for ~$16
1 article · October 10, 2026
- huggingface.coThe model that didn't exist, so you made it yourself
1 article · October 7, 2026
Topics and entities
Follow any of these and your For You feed starts watching them — no settings page required.