AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: AI And Watercolour Art: Training Models With TRL And OpenEnv on ThorstenMeyerAI.com

TL;DR

A developer has independently reproduced Surya Narreddi’s viral watercolour-generating AI model, using open tools like TRL and OpenEnv. All code, datasets, and models are now publicly available, enabling further research into aesthetic reinforcement learning.

An independent engineer has published a complete, open-source reproduction of Surya Narreddi’s viral watercolour AI model, using TRL and OpenEnv frameworks on Hugging Face infrastructure. This reproduction includes all datasets, training scripts, and trained models, enabling transparent evaluation of the model’s approach to aesthetic reinforcement learning. The original model, showcased in a widely viewed August video, was notable for its ability to generate watercolour-style paintings through code written by a language model, sparking significant interest in AI-generated art.

The reproduction leverages a reinforcement learning pipeline that trains a Qwen model with a reward system combining four terms: a compilation check, a length penalty, a style judge, and a preference model called HPSv3. The style judge, Qwen3-VL-30B-A3B-Instruct, assesses generated images against a set of four reference paintings, scored through a pairwise comparison process guided by human-like aesthetic preferences. All components, including datasets, environment scripts, and trained models, are openly available on Hugging Face, allowing researchers to replicate or extend the work.

This project aims to explore whether reinforcement learning can optimize AI models against aesthetic taste rather than factual correctness, a departure from typical RL applications that focus on verifiable answers. The model produces JavaScript code that creates watercolour paintings, making each brushstroke interpretable and editable—a feature that distinguishes it from pixel-based image generators. The open release addresses previous limitations, where only partial results or non-public artifacts existed, and invites further experimentation into aesthetic AI training methods.

At a glance
reportWhen: ongoing; the reproduction was published…
The developmentAn engineer has released an open, end-to-end reproduction of a watercolour AI model that gained viral attention in August, including datasets, training scripts, and models on Hugging Face.
At a glance
reportWhen: published after the 23 August viral vid…
The developmentA fully open reproduction of Surya Narreddi’s viral watercolour-painting coding model — including the RL environment, reference dataset, training scripts and trained models — has been published, built with TRL and OpenEnv and running entirely on Hugging Face infrastructure.

Implications for AI Art and Reinforcement Learning

This development is significant because it demonstrates that reinforcement learning can be applied to subjective aesthetic judgments, not just objective correctness. By openly sharing all artifacts, the project lowers barriers for researchers to investigate how models can be trained to favor artistic style and visual appeal, potentially influencing future AI art tools. It also raises questions about the role of human preferences in AI training and whether models can learn to produce more human-like, imperfect art that resonates on an emotional level. The approach challenges the prevailing trend of highly polished, perfect images from mainstream models, emphasizing the value of handcrafted, expressive outputs.

Amazon

watercolour digital art software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Evolution of AI Art and Reinforcement Learning Approaches

The project sits within a lineage of early AI art experiments, from DeepDream (2015) to neural network portraits by Mario Klingemann and datasets curated by artists like Anna Ridler. Historically, AI art has explored the medium’s creative potential, often through unsupervised or semi-supervised methods. Surya Narreddi’s original work, shared in August, marked a shift by combining language models with code-based painting tools, resulting in images that are both accessible and editable. The current reproduction builds on this foundation, emphasizing transparency and open science, which contrasts with proprietary or closed approaches often seen in commercial AI art tools.

Previous efforts in reinforcement learning for AI art have focused on objective rewards, such as passing tests or solving puzzles. Narreddi’s approach introduces a subjective, taste-based reward system, aligning more closely with human artistic preferences. This reflects a broader trend of integrating human feedback into AI training, but with a focus on aesthetic quality rather than correctness, representing a novel frontier in the field.

“This open reproduction allows the community to explore how reinforcement learning can optimize models for aesthetic appeal, a significant step beyond traditional, correctness-focused AI training.”

— Thorsten Meyer, AI researcher

Amazon

AI art generation tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Unconfirmed Aspects of Model Quality and Effectiveness

While all artifacts are publicly available, it remains unclear how closely the reproduction matches the original in terms of artistic quality. The comparative effectiveness of different reward mixes has not been quantitatively evaluated, and the influence of the specific dataset and environment setup on the results is still under investigation. Additionally, the broader applicability of this approach to other art styles or media has not yet been demonstrated, and the full technical report from Narreddi is pending publication, which could clarify some of these questions.

Amazon

watercolour painting digital brushes

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Future Research and Community Engagement Opportunities

Researchers and artists are encouraged to download and experiment with the shared artifacts to assess the model’s capabilities and limitations. The next steps include awaiting Narreddi’s full technical report for detailed methodology and results, as well as conducting comparative studies to determine which reward configurations yield the most aesthetically pleasing outputs. The open release also invites community-driven modifications, such as integrating new reference datasets or exploring alternative reward functions, to advance understanding of aesthetic reinforcement learning in AI art.

Amazon

artificial intelligence art kit

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

What does the open reproduction include?

The reproduction includes all datasets, training scripts, environment configurations, and trained models, all openly available on Hugging Face for community use and further development.

How does the model generate watercolour art?

The model writes JavaScript code that uses the p5.brush library to produce watercolour-style paintings, with each brushstroke decision being transparent and editable.

What is the significance of using reinforcement learning over taste?

This approach tests whether models can be optimized for subjective aesthetic preferences, moving beyond traditional correctness-based rewards and exploring new dimensions of AI creativity.

Are the results comparable to the original viral video?

It is not yet clear how closely the reproduction’s outputs match the original in quality, as a full technical evaluation and comparison are still pending.

What are the next steps for this research?

Future work includes analyzing the full technical report from Narreddi, experimenting with different reward configurations, and engaging the community in expanding or refining the approach.

Primary source: Hugging Face · via ThorstenMeyerAI.com

You May Also Like

The 9 Most Advanced AI Smartwatches To Buy In 2026

Discover the most advanced AI-powered smartwatches in 2026, including Apple, Samsung, Garmin, and budget options, based on ecosystem, battery, and features.

The High-End PC and Workstation Tax

Memory prices surge in 2026, making DIY PC building less cost-effective and impacting high-end workstations, with prices now comparable to prebuilt systems.

Alienware Surges In Global Coverage

Media coverage of Alienware has increased sharply, with 20 mentions in recent window—indicating rising interest, though the cause remains unconfirmed.

RHEO: Paint With Light

RHEO is a simple, calming app for iPhone, iPad, and Apple Vision Pro that transforms touch into flowing, beautiful light displays, emphasizing ease and calm.