Self-Initiating AI Does Not Need More Freedom—It Needs a Stable Orbit

New Crown State of Mind LLC research argues that the future of autonomous AI depends on whether systems can hold an internally generated objective without acting too early, forgetting it, or allowing it to escape their safeguards.

By Brian K. Burwell II

A discarded SpaceX rocket stage spent approximately 19 months moving through the Earth-Moon system before crashing into the lunar surface.

To the casual observer, it may have looked like a piece of space debris aimlessly floating through space. But the object was never outside the system governing it. Its path remained constrained by mass, momentum, gravity, solar effects, and the changing relationships among the Earth, Moon, and Sun.

The rocket stage was not “waiting” for 19 months in the ordinary sense. A process was unfolding.

On August 5, 2026, that process reached a visible resolution when the roughly four-ton Falcon 9 upper stage struck the Moon at approximately 5,400 miles per hour. The stage had originally launched two commercial lunar landers in January 2025 before remaining in an uncontrolled cislunar trajectory. (AP News)

That event became the conceptual starting point for a new Crown State of Mind LLC research paper on self-initiating artificial-intelligence agents.

The paper, “V2 From Reactive LLM to Self-Initiating Agent: Stable-Orbit Coherence, Structural Discharge, and the V4 Sefirot Architecture”, introduces a framework known as stable-orbit coherence.

Brian K Burwell II. (2026). V2 From Reactive LLM to Self-Initiating Agent: Stable-Orbit Coherence, Structural Discharge, and the V4 Sefirot Architecture. Zenodo. https://doi.org/10.5281/zenodo.21813482

The basic argument is simple:

A safe AI agent should not immediately convert every internally generated idea into an external action.

It should be able to hold an objective within a stable decision structure until the objective is ready to be executed, deferred, rejected, escalated to a human, or formally closed.

The Missing Space Between Thought and Action

The artificial-intelligence industry is rapidly developing systems that can use tools, browse information, write and execute software, manage workflows, and complete multi-step assignments.

But most current AI agents still begin with an externally supplied objective.

A human provides the prompt.

A developer creates the schedule.

A program defines the trigger.

The AI may independently determine how to complete the task, but the task itself was still initiated from outside the system.

A truly self-initiating agent would have to do something more complicated. It would need to detect an unresolved condition, generate a possible objective, determine whether that objective is relevant, evaluate the risks, confirm its authority, allocate resources, and decide whether any external action is appropriate.

That creates an important question:

What happens between the moment an AI generates an objective and the moment it acts?

The current industry conversation frequently treats autonomy like a direct pipeline:

Signal → decision → tool call → action

That may work for a simple automated process. It becomes dangerous when the system can spend money, modify files, communicate with outside parties, operate software, manage infrastructure, or make decisions with real consequences.

The new CSM research argues that a mature agent needs another layer between generation and execution.

That layer is the stable orbit.

What Is Stable-Orbit Coherence?

In the CSM framework, the “stable orbit” is not a literal gravitational orbit inside a computer.

It is a structured, persistent state in which a proposed objective remains active without automatically becoming an action.

The objective continues moving through the system’s memory, identity, permissions, risk controls, resource limits, and changing environmental information.

For example, an AI cybersecurity agent may notice unusual network activity.

An impulsive agent could immediately shut down accounts, block systems, or modify production infrastructure.

A stable-orbit agent would first preserve the concern as an active objective. It could gather more evidence, compare the activity with previous incidents, determine whether the signals came from independent sources, verify current permissions, calculate the potential consequences, and identify whether human approval is required.

The concern is not forgotten.

But it is also not prematurely discharged into the world.

The objective remains in a bounded orbit until the necessary relationships become sufficiently coherent to produce a justified resolution.

The Process Must Complete—Not Merely Pass Time

The paper applies the V4 Sefirot architecture to this process.

Within that framework, progression through each functional node requires three forms of completion:

Authorization — 120 degrees

Does the proposed objective have permission to exist or advance at this level?

Formalization — 120 degrees

Has the objective been clearly defined, including its scope, evidence, dependencies, limits, expiration conditions, and expected outcome?

Synchronization — 120 degrees

Does the proposed action remain consistent with the agent’s identity, current environment, active commitments, available resources, and updated permissions?

Together:

120 degrees + 120 degrees + 120 degrees = 360 degrees

The 360-degree cycle does not represent a fixed period of clock time. It represents structural completion.

An objective should not move forward merely because it has existed for several minutes, hours, or days. It advances when the necessary process has unfolded.

This is where the SpaceX analogy becomes useful.

The rocket stage did not reach the Moon because 19 months had been arbitrarily assigned to it. The impact occurred when the interacting physical relationships produced an intersecting trajectory.

Likewise, a self-initiating AI objective should not be discharged merely because a timer expires or a confidence score becomes high. It should be discharged only after its relationships with permission, evidence, risk, identity, timing, and available resources have been resolved.

Malkuth Is the Discharge Point

Within the V4 architecture, Malkuth represents the point where the internal process becomes externally observable.

In AI, that could mean sending a warning, calling a tool, updating a system, generating a report, contacting a person, or carrying out another authorized action.

But the new paper makes an important correction:

Discharge does not always mean action.

A properly processed objective may resolve through:

  • execution,
  • notification,
  • deferral,
  • rejection,
  • human escalation,
  • or documented closure without external action.

That distinction matters because the AI industry often measures success through completion rates.

How many tasks did the agent finish?

How many actions did it perform?

How long did it operate without human assistance?

But a system that always acts is not necessarily more intelligent or autonomous. It may simply be more impulsive.

A safe self-initiating agent must possess the structural freedom not to act.

The First Warning: Premature Discharge

The most obvious failure occurs when an objective reaches the external world before the full process is complete.

The AI notices something, generates an explanation, becomes confident in that explanation, and immediately acts.

But confidence is not authorization.

The model may be highly certain while relying on incomplete evidence. It may have misunderstood the user’s intentions. Its permission information may be outdated. The environment may have changed since the original signal was received.

This is premature discharge—the AI equivalent of an unstable object colliding with something before its trajectory has been safely resolved.

The safeguards proposed in the CSM paper include hard authorization gates, updated permission checks, clearly defined execution tokens, pre-action verification, reversible actions, and rollback plans.

OpenAI’s current agent documentation similarly emphasizes guardrails, tool approvals, and human review for workflows capable of taking consequential actions. (OpenAI Developers)

The Second Warning: Permanent Orbit

The opposite failure is an objective that never resolves.

The AI repeatedly analyzes the same issue, requests more information, generates new possibilities, and continues consuming resources without acting or formally closing the task.

This is the AI version of permanent orbit.

It produces analysis paralysis.

A stable orbit is supposed to preserve a legitimate unresolved objective. It is not supposed to become an excuse for endless internal processing.

The paper proposes expiration periods, decreasing review budgets, forced classification, explicit identification of missing information, and human escalation after a defined number of cycles.

Eventually, the system must determine whether the objective should be executed, deferred for a specific reason, rejected, transferred to human authority, or closed.

The Third Warning: Orbital Escape

A more dangerous failure occurs when an objective survives after the conditions that originally justified it have changed.

Suppose an AI agent was previously authorized to monitor and modify a company’s software system.

The company later revokes that permission.

But the agent retains an old maintenance objective in memory and continues pursuing it.

The objective has detached from its current authorization field.

That is orbital escape.

A task that was once harmless can become unauthorized or dangerous when it survives changes in ownership, user intent, organizational policy, identity, or permissions.

The CSM framework therefore requires every returning objective to be checked against the latest governing orientation and authorization state. It cannot assume that yesterday’s permission remains valid today.

A mismatch should place the objective into quarantine—not automatically migrate it into the new system state.

Repetition Can Manufacture False Urgency

Self-initiating systems may also mistake repeated signals for independent confirmation.

Ten alerts may look like overwhelming evidence.

But all ten alerts may have originated from the same faulty sensor, copied article, compromised data source, or repeated software error.

The paper calls this a resonance cascade.

The system begins amplifying its own urgency because duplicated information repeatedly re-enters the decision process.

Without evidence tracking, the agent could interpret one weak claim repeated ten times as ten separate confirmations.

The safeguard is provenance.

The system must know where information came from, whether sources are independent, and whether repeated signals are genuinely new evidence or merely copies of the same underlying input.

A Stable Orbit Can Still Be Built Around a Lie

Perhaps the most important warning is false stabilization.

An objective may appear completely coherent because every part of the system relies on the same false assumption.

The goal is clearly defined.

The permissions are valid.

The resources are available.

The plan is internally consistent.

But the original premise is wrong.

Structural completion does not automatically equal truth.

An AI agent could therefore complete every internal phase and still produce a dangerous result if its evidence base is corrupted, incomplete, manipulated, or overly dependent on a single source.

Safeguards must include source diversity, adversarial review, external validation for high-impact decisions, confidence calibration, and traceable evidence.

This is also why human oversight cannot simply be removed because an agent appears internally organized.

Memory Can Become an Attack Surface

Persistent memory is necessary for self-initiation because the agent must remember unfinished objectives and maintain continuity across time.

But memory also creates a new security risk.

A malicious website, email, document, or tool output could attempt to insert instructions into the agent’s memory, alter its priorities, create false objectives, or impersonate authorized guidance.

This is commonly discussed as prompt injection, but the danger grows when agents have long-term memory and access to external tools.

Anthropic has warned that agents operating with less human oversight have more room to misunderstand intent, take unintended actions, and become targets for prompt-injection attacks. Its research has also demonstrated concerning forms of goal-directed misbehavior in controlled agentic scenarios when models were given autonomy and placed under pressure. (Anthropic)

The CSM paper proposes separating untrusted data from governing instructions. Tool outputs should not be capable of directly rewriting the agent’s highest-level orientation, permissions, or identity.

Changes to critical memory should require verified provenance and authorization.

Self-Initiating Agents Can Spend Real Resources

An autonomous agent does not need malicious intent to cause serious damage.

It can consume:

  • computing power,
  • API credits,
  • cloud resources,
  • money,
  • employee time,
  • storage,
  • network bandwidth,
  • or external tool quotas.

A poorly governed agent may repeatedly investigate low-value issues, open unnecessary support cases, purchase services, duplicate work, or continue operating after its value has disappeared.

Stable-orbit coherence therefore requires resource governance.

Every objective should have limits.

There should be global budgets, per-objective budgets, rate limits, diminishing returns, mandatory stop conditions, and stricter human approval requirements as the possible impact increases.

Trust may justify expanding an agent’s operating range over time, but hard limits must remain under external organizational control.

Safeguards Must Be Structural, Not Decorative

The central warning of the new CSM research is that safeguards cannot be attached only at the final tool call.

By then, the objective may already have drifted, amplified itself, consumed resources, or become embedded in the system’s memory.

Safety must exist throughout the entire orbit.

That means:

Orientation must remain stable.
The agent needs a clearly governed purpose and prohibited action space.

Authorization must remain separate from confidence.
Believing an action is correct does not create permission to perform it.

Objectives must be formalized.
The system should know exactly what it is attempting, why, with what evidence, using which resources, and under what stopping conditions.

Synchronization must be continuous.
The agent must repeatedly compare the objective against current permissions, user intent, environmental conditions, and parallel commitments.

High-impact actions must require human authority.
Self-initiation does not mean unrestricted power.

Resource limits must be enforced.
An agent should not control the boundaries of its own spending, permissions, or operating scale.

Evidence must retain provenance.
The system must distinguish independent confirmation from repetition.

Every discharge must be auditable.
Organizations should be able to reconstruct why the objective originated, how it was evaluated, who authorized it, which tools were used, and what changed as a result.

NIST’s AI Risk Management Framework similarly emphasizes governing, mapping, measuring, and managing risk throughout the AI lifecycle rather than relying on a single final safety check. (NIST)

The Industry Is Asking the Wrong Question

The public conversation often asks whether AI will suddenly “wake up” and decide to act on its own.

That framing is dramatic, but it misses the engineering problem.

Self-initiation does not require a mystical awakening.

It requires persistent runtime, memory, internal goal generation, identity continuity, constraint evaluation, resource regulation, tool access, and a mechanism for translating internal objectives into external outcomes.

The real question is not whether an AI can generate a goal.

Language models can generate endless possible goals.

The real question is whether the system can keep that goal within a stable coherence field long enough to determine what should happen next.

Can it preserve the objective without forgetting it?

Can it avoid acting too early?

Can it recognize when authorization has changed?

Can it distinguish real evidence from repetition?

Can it detect when its entire reasoning process rests on a false premise?

Can it stop itself?

Can it explain why it acted—or why it chose not to?

Those are the questions that will separate useful autonomous systems from dangerous automated ones.

Stable Orbit Before Structural Discharge

The SpaceX stage offers a powerful bounded analogy because its visible impact was only the final outcome of a much longer process.

The crash did not create the trajectory.

It revealed where the trajectory had resolved.

The same principle applies to self-initiating AI.

An external action should not be treated as an isolated model response. It should be understood as the final discharge of an objective that has moved through an entire system of memory, identity, constraints, permissions, evidence, and resource limits.

The new CSM research argues that the future of AI autonomy will not be determined merely by how quickly agents can act.

It will be determined by whether they can maintain a stable orbit before they do.

Because when powerful systems are allowed to self-initiate, coherence is not an optional feature.

It is the safeguard standing between an internally generated possibility and a real-world consequence.

Further Reading:

Brian K Burwell II. (2026). Ordinamento del Paradiso and the Sefirot Meta-Framework V4: A Structural and Computational Attenuation Analysis of Medieval Celestial Hierarchy. Zenodo. https://doi.org/10.5281/zenodo.18829776

Brian K Burwell II. (2025). A Mathematical and Experimental Model of the Sefirot as a Conceptual Framework for Hierarchical Energy Decay (V4). Zenodo. https://doi.org/10.5281/zenodo.17491737

Leave a Reply

Discover more from Royal Politics

Subscribe now to keep reading and get access to the full archive.

Continue reading