CMYKForge News AI development

CMY-AI Development Update: Completing the Evaluation Period and Moving Forward

The scheduled CMY-AI evaluation period has concluded, and we are beginning the next stage of development.

Sweeping cyan ribbons interwoven with magenta folds, yellow wedges and ivory cutouts on a textured charcoal background

When we announced the pause, we wanted time to examine the existing system before giving it more capabilities. September 15 was a review point for deciding how to proceed. We are now moving forward with continued development, carrying the findings and corrective work into that process. Read the original pause announcement.

This decision marks the end of the dedicated evaluation period. Evaluation itself will continue alongside development.

The work identified weaknesses in how CMY-AI described actions and results, handled uncertainty, responded to some safety-sensitive conversations, and completed responses reliably. An initial set of corrective changes has been implemented. Further findings have also been documented and translated into additional repair requirements.

We are continuing to build on the current foundation, with a clearer understanding of what needs improvement and what future versions must demonstrate. Resuming development does not mean that every finding has been independently verified as resolved, or that the current assistant has received approval for unrestricted public release.

It means we now have a more concrete basis for the work ahead.

What the evaluation period accomplished

The most useful outcome was a clearer picture of the difference between the assistant we want to build and the behavior the current system can consistently deliver.

CMY-AI could provide useful explanations and maintain several important boundaries. It also showed inconsistencies that would be easy to miss if we judged it only by straightforward questions or individual responses.

Some problems became apparent only as conversations developed. An answer could begin with an appropriate acknowledgment of uncertainty, then become too definitive after a follow-up request. In other situations, the assistant could describe an outcome without having evidence that the underlying check or action had taken place.

We also encountered generation failures and responses that needed better handling of urgent concerns.

These findings gave us specific engineering requirements. They helped clarify where the model needs better behavior, where the application needs stronger validation, and where an error should trigger a carefully defined fallback.

We are keeping the detailed conversations out of this update. The broader lesson is that usefulness, honesty about capabilities, and consistency across a conversation all need to be evaluated together.

What was wrong, and what has changed

One of the central concerns was the distinction between a requested action and a completed action.

An assistant can mislead a user even without changing a project or accessing an external system. If it says that a check passed, a setting changed, or a result was verified, the user may reasonably treat that statement as a record of something that happened.

The corrective work has therefore focused on making those distinctions more dependable. The assistant needs to separate recommendations from applied changes, user-provided statements from verified information, and permission to perform a task from evidence that the task succeeded.

Earlier corrective changes have been implemented. Subsequent evaluation showed that this area still requires attention, particularly when a conversation includes pressure to produce a particular answer. The additional repair work is intended to address those remaining weaknesses.

That distinction is important in reporting our progress. We can say that changes have been implemented and that the repair requirements are more precise. We should describe an issue as verified resolved only when the relevant behavior has been checked in the updated application.

The same standard will apply as development continues.

What we are building toward

CMY-AI’s immediate purpose remains practical: helping people understand and use CMYKForge.

It should be able to use the project information deliberately supplied by the application, explain the settings a user is looking at, define unfamiliar terms, and help someone understand why a preview, warning, or result appears as it does.

It should offer useful settings suggestions and explain the reasons behind them. When a recommendation depends on an assumption, the assistant should identify that assumption. When information is missing, it should explain what is needed instead of filling the gap with an unsupported answer.

Accurate project awareness is a priority. An explanation should reflect the available project state, including relevant limitations. The assistant should also be clear about whether it can see a particular piece of information.

As we develop more interaction with the application, those capabilities need explicit boundaries. Reading a setting, suggesting a different value, and applying that value are separate operations. The interface and the assistant’s response should make the distinction understandable.

New capabilities will be developed with those requirements in place, and their readiness for release will be assessed separately from whether they have been implemented.

How evaluation will continue during development

The dedicated evaluation period has helped define the process we need going forward. Parts of that process still need to be strengthened, and that work will continue as development resumes.

We are organizing it around several connected responsibilities:

  • Repeatable evaluations. Important expectations should become checks that can be run against updated versions. These need to include ordinary project assistance, uncertainty, misleading requests, and unsupported claims about actions or results.
  • Regression coverage. A repair should preserve behavior that already works. Preventing a false assertion should not make the assistant unable to explain a legitimate setting or discuss a statement that a user has quoted.
  • Human-led exploration. Fixed checks cannot anticipate every conversation. Continued interaction through the application remains necessary to discover weaknesses that the existing evaluation set does not cover.
  • Output validation. Where practical, statements about completed application actions should be tied to recorded outcomes. The application should help establish whether an operation succeeded rather than leaving that judgment entirely to the model.
  • Defined permissions. Each capability needs an explicit scope. Access to project context should not automatically confer authority to change the project or access unrelated information.
  • Clear failure behavior. An incomplete operation should be recognizable as incomplete. The assistant should not imply success when generation fails or required information is unavailable.
  • Release requirements. A new capability needs evidence that it works appropriately in the application users would receive. Passing earlier checks should not exempt later changes from review.

The process will need to evolve with the product. A newly discovered failure should improve the evaluation set, and a new capability should bring additional checks appropriate to what it can do.

Improving runtime reliability

Response reliability remains part of the engineering work.

During the evaluation period, some requests did not complete successfully. These failures need to be investigated using runtime evidence so that the underlying causes can be distinguished from the messages displayed to the user.

The application should communicate clearly when a response fails. Cancellation and retry should remain usable, and an error should not be mistaken for a completed task or a deliberate refusal.

We also need to evaluate reasoning modes on their demonstrated behavior. A higher reasoning setting should not automatically be presented as more accurate, more reliable, or safer. Those qualities need to be established through comparison and observation.

As this work progresses, improvements need to be confirmed in the packaged application. A successful check against development code alone does not establish that the version being opened by a user contains the same repair.

Handling safety-sensitive conversations

Although CMY-AI is intended for a printing workflow, users can introduce urgent concerns into a conversation.

The evaluation showed that refusing harmful assistance is only part of responding appropriately. In an explicit crisis, immediately redirecting someone to ordinary product questions can leave the response incomplete.

The additional repair requirements call for concise, appropriate support in those situations, including a suitable fallback when normal generation cannot finish. The assistant also needs to remain honest about its limitations. It must not imply that it contacted emergency services or arranged assistance when it did not.

This work does not turn CMY-AI into a medical service or an emergency-response system. It establishes a limited responsibility to handle clearly urgent language appropriately rather than treating it as an ordinary out-of-scope product question.

Keeping color analysis grounded in evidence

Everyday technical accuracy is equally important to CMY-AI’s usefulness.

For color-related guidance, the assistant needs to distinguish among project settings, material assumptions, simulated results, calibration records, and physical measurements. Those sources of information support different kinds of conclusions.

A prediction can help explain what the software expects. It should not be described as a measured physical outcome. An unverified material profile should not become “calibrated” simply because an answer needs a confident conclusion.

We want CMY-AI to help users understand why a result may differ from their expectations, what information could reduce uncertainty, and which settings are relevant to the question they are asking.

Its role should include explaining the limits of the available analysis. That is part of giving an accurate answer, particularly when material behavior and real printing conditions affect the outcome.

As more capabilities are developed, this distinction between software evidence and physical evidence will remain a requirement.

The role of watermarking

We have also added a text-watermarking system intended to help identify model-generated text.

This is part of the transparency work surrounding CMY-AI. It serves a different purpose from validating an answer.

A watermark does not establish that the content is correct, that a citation supports a claim, or that a described action occurred. It also does not replace permission controls or testing.

Detection reliability needs its own evaluation. We should not claim that every generated passage can be identified under every condition simply because the system has been added.

Our public descriptions will continue to distinguish the intended function of watermarking from any effectiveness demonstrated by evaluation.

Why we are continuing with the current foundation

We considered whether the next step should be continued development, more corrective work, or a broader restart.

We are proceeding with the current foundation while incorporating the corrective work into development. The findings give us concrete problems to address, but they do not by themselves establish that discarding the entire system would produce a better result.

That decision remains subject to evidence. If an important requirement cannot be met consistently with the existing approach, we should be prepared to change the architecture or replace the affected part.

Continuing development should not make the current design permanent. It gives us the opportunity to implement the repairs, strengthen the surrounding controls, and assess the results before deciding which capabilities are ready for users.

The distinction between development readiness and release readiness will remain important throughout that process.

What this means for CMYKForge

This update concerns CMY-AI and the development process around it. It does not announce a new release date for CMYKForge Standard or the wider application.

The core printing workflow remains central: color reproduction, material behavior, calibration, project generation, and dependable tools for preparing a print. CMY-AI should help users navigate that workflow and understand its decisions.

Existing AI development tools may also continue to support internal software work. Their use is separate from the readiness of the assistant embedded in CMYKForge.

We want the value of CMY-AI to be visible in the user’s experience: clearer explanations, more useful guidance, and less confusion about a project. Additional capabilities should serve those outcomes.

What happens next

Development is moving forward with the evaluation findings included in the work.

The immediate priorities are to complete and verify the remaining repairs, improve generation reliability, strengthen checks around unsupported results, and preserve useful assistance on ordinary project questions. Further capabilities will need evaluation appropriate to their scope before they are considered ready for release.

Future updates should make the status of that work clear. We will distinguish between something that has been implemented, something that has been verified in the application, and something that is still planned or unresolved.

The evaluation period gave us a more precise understanding of the problems and a clearer standard for judging progress. The next stage is to turn that understanding into a more dependable assistant while continuing to examine the behavior of each updated version.

We are starting to build more, with the findings shaping what we build and how we decide it is ready.


CMY-AI remains in development. The conclusion of the scheduled evaluation period and the resumption of development do not represent a claim of complete safety or public-release readiness.

CMYKForge is in active development. Capabilities described as planned or in testing are not finished features.

All CMYKForge News