CMYKForge News AI development

CMYKForge Is Pausing CMY-AI Development to Build a Stronger Safety and Evaluation Framework

Abstract CMYKForge graphic of a glowing cyan, magenta, yellow and white ring on a dark field bordered by diagonal color bands

As CMYKForge moves closer to the public release of Standard, we are making a deliberate change to the development of CMY-AI.

Effective August 20th, we have paused further training and capability development of CMY-AI through September 15th.

This is not a pause in CMYKForge development. Work on CMYKForge Standard and the rest of the application will continue. We will also continue using existing AI development tools where they help us build and test CMYKForge. The pause applies specifically to the continued development of CMY-AI, our AI system designed to operate within CMYKForge.

Testing of CMY-AI is also not stopping. In fact, during this period, we expect testing to continue at least as extensively as before and, in some areas, increase. The purpose of the pause is relatively simple: before making CMY-AI more capable, we want better ways of measuring and understanding the capabilities it already has.

Why we are pausing development

CMY-AI is not intended to be only a chatbot sitting beside CMYKForge. The system is being designed to understand CMYKForge, answer questions, analyze project inputs and previews, modify settings, and interact with significant portions of the application. That potential makes CMY-AI useful. It also means we believe it deserves more scrutiny than an ordinary conversational assistant. As development has progressed, we have been able to manually test the model, experiment with different situations, examine its responses, and look for behaviors we do not want. That work will continue. But there is an important limitation to this approach.

The behavior we personally manage to find is not necessarily the full range of behavior a model is capable of producing.

Testing a model repeatedly and seeing good results provides useful information. It does not prove that every important boundary is working correctly. And continuing to add capabilities while the system used to evaluate those capabilities is still relatively manual can make that problem harder. We have therefore decided that this is the right point to stop expanding CMY-AI temporarily and put more attention into evaluating what already exists.

This is a development freeze, not a testing freeze

During the pause, we do not plan to continue adding new training data or intentionally expanding CMY-AI’s capabilities.

We do plan to continue interacting with the existing model.

That distinction is important.

Freezing the model gives us a more stable target to evaluate. Instead of testing something that is continuously changing underneath us, we can spend more time examining a known version of the system. We can repeat tests. We can compare results. We can document failures. We can look for inconsistencies. And when we find something that CMY-AI handles incorrectly, we can turn that discovery into a test that future versions should have to pass. The objective is not to spend several weeks doing nothing with CMY-AI. The objective is to spend that time learning more about it before making it more capable.

Manual testing is not enough

Manual testing has been an important part of CMY-AI’s development and will remain important. A person can notice behavior that an automated test might miss. Unusual conversations can expose problems that were never anticipated when a test suite was written. But manual testing has another weakness: you can only test what you think to try. CMY-AI needs a more systematic evaluation process alongside that human testing. After CMYKForge Standard reaches its planned public release point, we intend to dedicate an initial period of at least two weeks specifically to strengthening that infrastructure before normal CMY-AI development resumes. That work is planned to include several areas.

Automated evaluations

We want a repeatable collection of tests that can be run against CMY-AI as it changes. Instead of relying entirely on whether a developer remembers to test a particular situation, important behaviors can become part of a permanent evaluation set. As problems are discovered, new tests can be added. Over time, that should create a growing record of behaviors that CMY-AI is expected to handle correctly.

Regression testing

Fixing one problem should not silently recreate another. Future CMY-AI versions should be tested against behaviors previous versions were already expected to handle correctly. If an update improves one capability but causes an important existing test to fail, that should be visible before the model reaches users.

Adversarial testing

Normal questions are only one part of evaluating an AI system. We also intend to deliberately test CMY-AI with unusual, conflicting, misleading, unexpected, and boundary-pushing inputs. The purpose is not to assume that customers will deliberately try to break CMY-AI. It is to recognize that software should be evaluated beyond the easiest conditions.

Output validation

CMY-AI can potentially interact with important parts of a CMYKForge project. That means an AI-generated decision should not automatically be treated as correct simply because the model produced it confidently. Where appropriate, deterministic parts of CMYKForge should validate model outputs before those outputs are allowed to affect a project. AI reasoning and conventional software validation can serve different purposes, and we believe both have a role.

Permission and action boundaries

A model capable of interacting with an application should not automatically have unrestricted authority over everything the application can do. Part of this work will involve examining what CMY-AI should be able to control, under what circumstances it should be able to control it, and where additional validation or user involvement should be required.

The question is not simply:

Can CMY-AI perform this action?

It is also:

Should CMY-AI be allowed to perform this action automatically?

Those are different questions.

Failure behavior

A useful AI system should also know what happens when it cannot reliably complete something. Sometimes the correct behavior may be asking the user for more information. Sometimes it may be handing the task back to a deterministic part of CMYKForge. Sometimes it may mean refusing to perform an action because required information is missing or the requested operation falls outside its permitted boundaries. We do not want confidence alone to determine whether an action should happen.

Monitoring

Testing before release is important, but it cannot represent every situation an AI system may encounter. We therefore intend for evaluation to continue throughout CMY-AI’s development rather than treating safety work as a one-time phase that eventually becomes “finished.” Problems discovered during development and future testing should feed back into the evaluation system. A failure found once should become something we know how to look for again.

Release gates

Eventually, we want CMY-AI development to operate with clearer requirements for what must be tested before an updated model is considered ready. A new version being newer will not automatically make it better. An updated model should demonstrate that it still meets the standards expected of the previous version while also being evaluated for whatever new capabilities have been introduced.

September 15 is a review point, not an automatic restart

We have previously said that the current development pause is planned through September 15. We want to clarify what that date means.

September 15 is not a promise that CMY-AI development automatically resumes that day.

It is the point at which we intend to formally evaluate what we have learned about the current model and decide what happens next. There are multiple possible outcomes. If testing gives us sufficient confidence in the existing foundation and we have the evaluation infrastructure needed to continue responsibly, we can move toward resuming development under those stronger controls. If significant problems remain, the pause can continue. And if our analysis shows that the existing CMY-AI foundation does not meet the standards we want for the product, we are prepared to discard the current training work and start that process again with stricter requirements. That would cost development time. We believe that would still be preferable to continuing simply because significant time had already been invested in the current version. Past development effort should not become a reason to ship something we no longer believe is the right foundation.

Two weeks will not make an AI system “safe”

We also want to avoid creating the wrong impression about the period following Standard’s release. We are not claiming that two weeks of work can prove that an AI system is completely safe.

It cannot.

The planned two-week period is intended to establish a stronger initial safety, evaluation, and monitoring framework around future CMY-AI development. That infrastructure is the beginning of an ongoing process, not the end of one. As CMY-AI changes, its evaluations should change with it. New capabilities can create new failure modes. New problems can require new tests. A system that passed an evaluation previously should not receive permanent approval regardless of how much it changes afterward. Our intention is for testing to become part of CMY-AI development itself.

Safety and quality are not the same thing

There is another distinction we think is important. Not every bad AI output is necessarily a safety problem. If CMY-AI recommends an inappropriate setting, misunderstands part of a project, gives incorrect information about CMYKForge, or produces a poor recommendation, that may be primarily a reliability or quality failure. Those failures still matter.

CMYKForge’s product priorities remain:

  • Print quality
  • Color accuracy
  • Reliability
  • Ease of use
  • Performance

CMY-AI should ultimately be held to standards that support those priorities. Our evaluations therefore should not focus only on traditional AI-safety questions. They also need to examine whether CMY-AI understands CMYKForge accurately, follows instructions reliably, handles uncertainty appropriately, respects its permissions, protects user expectations, and actually improves the workflow it is supposed to assist. An AI feature that is technically impressive but unreliable is not automatically a good CMYKForge feature.

CMYKForge Standard is not being paused

We want to make this particularly clear.

Development of CMYKForge is continuing.

This decision applies to CMY-AI development specifically.

CMYKForge Standard remains the immediate priority as we work toward public availability. The underlying purpose of CMYKForge is full-color FDM image reproduction. The core product is being built around color science, material behavior, calibration, project generation, and reliable printing workflows. CMYKForge does not need to become an AI-first product in order to accomplish that mission. AI may eventually make parts of the experience substantially better. Where it does, we want to use it. But we do not believe that every part of CMYKForge needs AI, nor do we believe that CMY-AI should become more capable simply because making it more capable is technically possible.

This does not mean CMYKForge is abandoning AI

We are not abandoning CMY-AI.

We are also not stopping the use of existing AI development tools in the broader development of CMYKForge.

Those are separate things. AI tools can continue assisting with software development and other appropriate internal work while the model that will become CMY-AI remains frozen for evaluation. The purpose of this decision is not to reject AI. It is to establish a higher standard for an AI system that could eventually interact directly with CMYKForge users and their projects.

Our position on the wider AI discussion

There is a much larger discussion happening about how quickly increasingly capable AI systems should be developed and deployed. CMYKForge is not an AI research laboratory, and we do not pretend to have the authority to dictate how the entire AI industry should approach that question. What we can control is our own product. For us, responsible AI development should mean being willing to stop adding capabilities when our ability to evaluate them needs to catch up. It should mean being willing to discover that an approach is not good enough. It should mean being willing to throw work away rather than allowing the amount of time already invested in it to determine what customers eventually receive. And it should mean recognizing the limits of what testing can prove.

What happens next

Between now and September 15, CMY-AI will remain a stable target for continued testing rather than continued capability expansion.

We will use that period to learn more about its existing behavior and document what needs to become part of the formal evaluation process.

September 15 will serve as a formal review point.

From there, the results—not the calendar alone—will determine the next step. If the foundation is strong enough, we can move forward with better safeguards and more systematic monitoring. If it needs more work, we will give it more work. If the foundation itself is wrong, we are prepared to rebuild it. CMY-AI has the potential to become an important part of CMYKForge Professional. That is exactly why we do not think we should rush it.

Before we make CMY-AI more capable, we want to become better at understanding the capabilities it already has.

CMYKForge is in active development. Capabilities described as planned or in testing are not finished features.

All CMYKForge News