🔍 Read the full analysis: Why Build An AI Model From Scratch? on ThorstenMeyerAI.com
Get the latest gadgets delivered free with Prime
- Fast, free delivery on millions of items
- Prime Video, Amazon Music and more included
- Member-only deals all year
TL;DR
A Hugging Face contributor says its ML-intern agent helped build and publish seven custom models over several days, including a small prompt rewriter and a citrus-diagnosis model. The contributor reported low compute costs and task-specific performance gains, but the figures are self-reported and have not been independently replicated.
A Hugging Face contributor says the platform’s ML-intern agent helped build and publish seven custom models over several days, including a 0.8-billion-parameter prompt rewriter and a model for identifying citrus problems from photos. The examples offer a practical account of agent-assisted model development, as described in the original analysis, but the contributor’s reported results and costs are not independently verified.
The contributor said the work began with a request for a smaller version of the prompt rewriter included with Qwen-Image 2.1, an example of the kinds of specialized tools discussed in AI model selection. According to the account, the original model has 9 billion parameters, requires about 20 GB of memory and can generate thousands of tokens before producing a paragraph. The contributor reported that the resulting 0.8B model produced valid output 99.7% of the time and used about one-quarter as many tokens as its teacher model. Compute, including labeling 8,797 example requests with the larger model, cost about $16, the contributor said.
Another project fine-tuned Qwen3.5-2B to identify citrus pests, diseases and nutritional deficiencies in images, illustrating how custom models can be tailored to specific tasks, much like the models covered in this look at capable AI models. The contributor said the training set contained 3,017 annotated images across 21 categories. On 335 test photos, the base model reportedly selected the correct problem 14.9% of the time, compared with 52.8% for the fine-tuned model after two training epochs on one A10G GPU. The reported compute cost was about $1.90.
The account also describes a character-generation LoRA trained on 84 captioned drawings and a camera-angle LoRA for Qwen-Image 2.1. For the latter, the agent generated 24,722 transparent images of scanned household objects across 24 angles. The contributor said training took about 90 minutes on one A100 and the project’s compute cost was about $16, including failed jobs that had to be resubmitted. Model cards and evaluations were published on Hugging Face, according to the account; detailed descriptions were supplied for only some of the seven projects.
Lowering the Barrier to Custom Models
The examples matter because they show how an agent could reduce the coordination work involved in making a specialized model. The contributor described asking ML-intern to plan a project, run an initial test, train a model, evaluate it and publish the results using Hugging Face hardware. For developers with a narrow need, that could make small-scale experimentation more accessible than managing every step manually.
But the reported dollar figures cover compute costs, not necessarily the full cost of development. Preparing and checking data, writing prompts, reviewing model outputs and deciding whether the evaluation is sound also take time. The figures are examples from one contributor’s projects, not a price guarantee or evidence that similar results are typical for other users.
The citrus example also illustrates why a baseline matters: its reported score is compared with the untuned model on the same test set. By contrast, the account says later character LoRA checkpoints began affecting prompts unrelated to the intended character. That observation points to a practical risk in customization: a model can learn a desired style too broadly, so training gains need to be checked against unintended effects.
As an affiliate, we earn on qualifying purchases.
How the Agent Ran Each Project
The contributor said each project began as a message in HuggingChat with ML-intern enabled. The agent proposed a plan and requested spending approval before paid work. If a prompt did not include a budget, it reportedly offered options and asked the user to choose. The contributor said the workflow then moved through a small test run, training, evaluation and publication.
Over the course of the projects, the contributor’s prompts grew from about 450 words for the first effort to nearly 2,000 for the sixth. They included information about the dataset, base model and training script, as well as requests to establish a baseline, run a smoke test and follow a spending cap. The contributor said all seven prompts are available in a public GitHub repository. This account documents one person’s workflow; it is not an independent assessment of ML-intern or a comparison with other model-building methods.
“Also report the base model’s zero-shot score on the same metric before training so we can see the gain.”
— The Hugging Face contributor, describing a prompt used for model evaluation
As an affiliate, we earn on qualifying purchases.
What the Report Cannot Establish
The performance figures, project descriptions and cost estimates are self-reported. The supplied account does not provide independent replication, complete evaluation protocols for every model or comparable results from other users. It also does not give detailed descriptions of all seven projects.
Some measurement details remain unclear. The account does not explain how the 99.7% valid-output rate was calculated, whether test images were independently reviewed, or how data quality was checked across projects. The reported results may also change on images or prompts outside the test sets, and the cost figures do not account for all time and possible expenses. These gaps limit how confidently readers can generalize from the examples.
machine learning model development kit
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Published Models Need Wider Testing
The contributor says the models, evaluations and prompts are available through Hugging Face and GitHub, where readers can inspect the published work. Wider comparisons across tasks, datasets, hardware and users would help show whether the reported costs and performance changes are reproducible.
The contributor’s stated process recommends recording a baseline before training, running a small test and checking that saved model weights changed, and setting a spending limit in advance. Whether those steps provide adequate checks will depend on the task and the quality of the data. No independent follow-up results or broader evaluation were included in the supplied account.
As an affiliate, we earn on qualifying purchases.
Key Questions
What did the Hugging Face contributor report?
The contributor said the ML-intern agent helped build and publish seven custom models over several days. The account gives detailed examples for a prompt rewriter, a citrus image model and two LoRA projects, but not all seven.
How much did the projects cost?
The contributor reported about $1.90 in compute for the citrus model and about $16 for the prompt rewriter and camera-angle LoRA project. These figures refer to reported compute spending, not a complete accounting of development costs.
Were the reported model results independently verified?
No independent replication is described in the supplied account. The reported performance figures and project details come from the contributor, and evaluation methods are not fully specified for every model.
Why build a model from scratch or customize one?
In these examples, the aim was to adapt or create smaller, task-specific models—for instance, a prompt rewriter intended to use fewer resources or a model trained to classify citrus problems from images. The account describes fine-tuning and LoRA projects as well as a smaller model; it does not establish that every project involved training a model entirely from random initialization.
Primary source: Hugging Face · via ThorstenMeyerAI.com
Halloween Picks
halloween
As an affiliate, we earn on qualifying purchases.
