Traditional Chinese version: [[content/posts/operations-hospital-collaboration-guide]]

1. Purpose

This guide is intended for the TAMS Backend, Hospital Integration, RMF/Robot Integration, QA, and Operations teams. It describes the actual boundaries of the current Modular Monolith, how Operations and Hospital communicate, and the rules that all contributors must follow when changing the codebase.

The system is still deployed as one NestJS process, one npm package, and one versioned release. Modules communicate through NestJS dependency injection and in-process method calls. There are no internal HTTP calls, message brokers, or additional worker frameworks between modules.

The primary design principle is:

Hospital decides what the hospital workflow requires. Operations guarantees how robot tasks are executed safely. Adapters handle how the system communicates with external technologies.

2. Quick Placement Guide

RequirementOwner/locationExamples
General task, robot, or facility rulessrc/operations/<feature>Task cancellation, dispatch, robot availability, charging-station lookup
Background workflows spanning multiple Operations featuressrc/operations/workflowsQueue Worker, Robot Keeper
CSH/CMP/HIS or hospital-specific workflowssrc/csh-hospitalEmergency return to pharmacy, employee sync, medication lookup
HTTP validation and response mappingsrc/product-api, src/csh-hospital/api, or the existing controller layerDTOs, guards, Swagger decorators
MongoDB, Redis, RMF, Axios, ROS, or Socket.IOsrc/adapters or the existing infrastructure layerRepositories, RMF adapters, robot-control adapters
Authentication, JWT, RBAC, sessions, and auditsrc/accessActor, permissions, user management
Pure types shared by multiple top-level modulessrc/domainTypes that do not belong to one Operations feature

If a requirement contains a hospital-name condition such as if (hospital === 'CSH'), do not add it directly to Operations. First determine whether configuration can express the difference. Introduce a small, explicit cross-module contract only when the behavior is genuinely different.

3. System Architecture

flowchart LR
  ProductAPI[Product API controllers]
  HospitalAPI[Hospital API controllers]
  HospitalWF[CSH Hospital workflows]

  subgraph Operations[Operations Module]
    Task[Task feature]
    Robot[Robot feature]
    Fleet[Fleet feature]
    Facility[Facility feature]
    Settings[Settings feature]
    CrossWF[Queue Worker / Robot Keeper]
  end

  SharedPorts[Operations shared ports]
  LocalPorts[Feature-local outbound ports]
  Adapters[Mongo / Redis / RMF / Robot control / WebSocket adapters]

  ProductAPI -->|feature public entry| Task
  ProductAPI -->|feature public entry| Robot
  ProductAPI -->|feature public entry| Fleet
  ProductAPI -->|feature public entry| Facility
  ProductAPI -->|feature public entry| Settings

  HospitalAPI --> HospitalWF
  ProductAPI -->|medicine-cart endpoint| HospitalWF
  HospitalWF -->|TaskProgressionService public entry| Task
  HospitalWF -->|EmergencyOperationsPort| SharedPorts
  SharedPorts -->|implemented by emergency facade| CrossWF
  HospitalWF -->|SiteOperationsConfig implementation| SharedPorts

  Task --> LocalPorts
  Robot --> LocalPorts
  Fleet --> LocalPorts
  Facility --> LocalPorts
  Settings --> LocalPorts
  CrossWF --> Task
  CrossWF --> Robot
  LocalPorts -->|implemented by| Adapters
  SharedPorts -->|events/config adapters| Adapters

3.1 Dependency Direction

The normal dependency direction is:

API / Hospital workflow
        ↓
Operations application
        ↓
Domain + Ports
        ↑
Adapters implement Ports

Operations must not import a concrete CSH Hospital implementation. Hospital obtains general task capabilities through the task feature’s public entry and executes complete cross-feature Operations use cases through EmergencyOperationsPort. Operations returns neutral outcomes and never calls back into Hospital.

4. Feature-First Operations Structure

src/operations/
├── task/
│   ├── domain/
│   ├── application/
│   │   └── ports/
│   ├── public/
│   └── task.module.ts
├── robot/
│   ├── domain/
│   ├── application/
│   │   └── ports/
│   ├── public/
│   └── robot.module.ts
├── fleet/
├── facility/
├── settings/
├── workflows/
├── ports/
└── operations.module.ts

4.1 Responsibilities Within a Feature

domain

  • Contains business entities, enums, records, value objects, and pure rules.
  • Must not depend on NestJS, Swagger, MongoDB/Mongoose, Redis, Axios, ROSLIB, or Socket.IO.
  • Must not contain HTTP requests, Axios responses, Mongoose documents, or raw RMF payloads.
  • IDs crossing domain, application, and port boundaries must be strings.

application

  • Executes a concrete use case and coordinates domain logic with ports.
  • May use NestJS dependency injection, but must not import concrete adapters, API DTOs, or Hospital implementations.
  • Must not construct Axios, RMF, or Redis wire payloads directly.

application/ports

  • Defines the external capabilities required by the owning feature, such as repositories, runtime-state stores, RMF commands, or robot control.
  • Ports should be designed around use-case needs rather than generic CRUD repositories.
  • Port commands and results must be transport-neutral.

public

  • Provides the public import entry used by APIs, Hospital, CLIs, and inbound adapters.
  • The current public/index.ts is a controlled export barrel, not an additional proxy or separate service instance.
  • Before exporting a new service, verify that an external module genuinely needs it. Do not re-export every application service by default.

*.module.ts

  • Registers and exports providers owned by the feature.
  • Feature modules do not contain HTTP controllers.
  • OperationsModule only composes feature modules and cross-feature workflows.

4.2 Feature Ownership

FeatureResponsibilitiesMain public capabilities
TaskTask history, task libraries, queue intent, dispatch, cancellation, continuationTaskApplicationService, TaskRequestService, TaskCancelService
RobotRuntime state, availability, navigation, snapshots, videoRobotApplicationService, RobotVideoService, UpdateRobotStateService
FleetFleet CRUD, fleet commands, Fleet Manager discoveryFleetApplicationService, FleetService
FacilityMaps, doors/lifts, devices, charging lookupDeviceApplicationService, MapFloorService
SettingsSystem settingsSystemSettingsApplicationService

TaskQueueWorker and RobotKeeperWorker depend on several features. They therefore live in operations/workflows and are registered by OperationsModule.

5. CSH Hospital Module

src/csh-hospital/
├── api/           Hospital HTTP controllers
├── config/        CSH/site configuration and typed access
├── dto/           Hospital transport DTOs
├── integrations/  CSH/CMP/HIS HTTP integrations
├── ports/         Hospital-owned persistence/integration ports
├── workflows/     Hospital business workflows
└── csh-hospital.module.ts

The Hospital module is responsible for:

  • Interpreting CMP/HIS/CSH requests and event semantics.
  • Employee, medication, and hospital-data integrations.
  • Deciding which hospital workflow an emergency requires.
  • Using Operations capabilities to cancel, create, pause, or continue tasks.
  • Providing site and hospital configuration without forcing Operations to inspect hospital names.

The Hospital module must not:

  • Build RMF payloads or call Open-RMF directly.
  • Operate on Redis keys or MongoDB collections directly.
  • Reimplement queue leasing, robot availability, or dispatch idempotency.
  • Bypass Operations cancellation and dispatch safety rules.

6. Communication Between Operations and Hospital

6.1 Hospital Calls Operations

A Hospital controller first calls a Hospital workflow. The workflow then uses task and robot capabilities provided by Operations.

For an emergency event:

  1. POST /cmp/messages/hospital receives ReturnToErPharmacy.
  2. HospitalMissionControlService.requestForActiveTasks() first persists fleet-wide taskAcceptancePausedAt and a new taskAcceptancePauseVersion in MongoDB.
  3. It obtains active task contexts through EmergencyOperationsPort. If MongoDB or Redis cannot be read, the flow fails closed and the persisted fleet pause remains active.
  4. It broadcasts EMERGENCY_STARTED to every robot, not only active robots.
  5. For every active task, it decides whether:
    • The robot is moving: cancel the original task and create a return-to-pharmacy task.
    • The robot is already at a non-pharmacy station: mark a return request and wait for HMI continue.
  6. A return failure for one task does not block other tasks; that item returns EMERGENCY_RETURN_REQUIRES_RECOVERY.
  7. It returns notifiedRobots and the processing result for each task. emergency.changed is published after the Mongo pause succeeds.

HospitalMissionControlService injects only EmergencyOperationsPort and Hospital configuration. Repositories, Redis runtime state, HMI control, cancellation safety, and queue creation stay encapsulated in EmergencyOperationsService, so Hospital does not know how Operations data is manipulated.

6.2 Medicine-Cart Continue Orchestration

When the HMI continues a task, Product API calls MedicineCartTaskWorkflow. This Hospital workflow then calls Operations TaskProgressionService:

  • For a normal task, Operations queues the next station and returns a continued outcome.
  • When an unhandled returnToErPharmacyRequestedAt exists, Operations returns emergency-return-required with a transport-neutral task context.
  • Hospital chooses the return station, invokes EmergencyOperationsPort.cancelAndQueueReturn(), and maps the result to the unchanged HTTP response shape.

The Operations outcome contains no Hospital class, HTTP DTO, or CSH configuration:

type TaskContinuationOutcome =
  | { kind: 'continued'; result: ContinueTaskResponse }
  | { kind: 'emergency-return-required'; task: EmergencyContinuationContext };

There is no Operations-to-Hospital reverse dependency and no Hospital provider token bound back into TaskModule. TaskProgressionService is registered normally by TaskModule; MedicineCartTaskWorkflow is registered by CshHospitalModule and exported to the controller.

6.3 Configuration Flow

Operations reads the following values through SiteOperationsConfig:

  • siteName
  • returnHomeStation
  • taskSimulation

CshHospitalConfig implements this contract. AdaptersModule binds it to SITE_OPERATIONS_CONFIG using useExisting.

To add site-level configuration required by Operations:

  1. Add a transport-neutral field to SiteOperationsConfig.
  2. Implement it in the Hospital configuration while preserving backward-compatible defaults.
  3. Add configuration unit tests and Operations use-case tests.
  4. Update .env.example and deployment documentation.

Do not inject ConfigService into Operations to read hospital-specific environment variables. Do not read process.env directly while loading application modules.

6.4 Business Event Publication

Operations and Hospital publish these events through OperationsEventPublisher:

  • robot.state.changed
  • task.state.changed
  • fleet.state.changed
  • emergency.changed

The current implementation is consumed by a WebSocket adapter. This publisher supports dashboard and event notifications; it is not a durable event bus. It must not replace MongoDB state changes or queue transactions that are required to succeed.

7. Emergency Return-to-Pharmacy Flow and Invariants

sequenceDiagram
  participant CMP
  participant HospitalAPI
  participant HospitalWF as HospitalMissionControl
  participant Ops as EmergencyOperationsPort
  participant Mongo
  participant HMI as RobotControlPort
  participant Queue as Task Queue Worker
  participant RMF as FleetCommandPort
  participant ReturnOp as EmergencyReturns

  CMP->>HospitalAPI: ReturnToErPharmacy
  HospitalAPI->>HospitalWF: requestForActiveTasks()
  HospitalWF->>Ops: pauseFleetTaskAcceptance()
  Ops->>Mongo: persist pause timestamp + version
  HospitalWF->>Ops: listActiveTaskContexts()
  Ops->>Mongo: load active tasks and ownership
  Ops->>Ops: read normalized Redis runtime state
  HospitalWF->>Ops: broadcastEmergency(EMERGENCY_STARTED)
  Ops->>HMI: notify every robot
  loop each active task
    alt robot is moving
      HospitalWF->>Ops: cancelAndQueueReturn()
      Ops->>ReturnOp: claim by original task id
      Ops->>ReturnOp: phase=CANCELLATION_REQUESTED
      Ops->>RMF: cancel booking
      RMF-->>Ops: explicitly accepted
      Ops->>Mongo: cancel original task locally
      Ops->>ReturnOp: phase=ORIGINAL_CANCELED
      Ops->>Mongo: create return task + queue intent
      Ops->>ReturnOp: COMPLETED / RETURN_QUEUED
    else arrived at non-pharmacy station
      HospitalWF->>Ops: markReturnRequested()
      Ops->>Mongo: persist return request + emergency version
    end
  end
  HospitalWF-->>CMP: affectedTasks + notifiedRobots
  Queue->>RMF: dispatch return task safely

Changes to this flow must preserve all of the following:

  • A fleet pause blocks ordinary tasks and Robot Keeper auto-charging tasks.
  • Manual cancellation remains rejected while the active robot is emergency-paused.
  • Only the internal emergency-return workflow may use allowDuringEmergency: true.
  • Emergency return must also use requireFleetConfirmation: true. Local queue deletion, task-history cancellation, and return creation happen only after RMF explicitly accepts cancellation.
  • A rejected or unknown RMF cancellation must not be retried automatically and must not create a return task. The durable operation remains available for manual reconciliation.
  • The emergency-return cancellation uses notifyAmr: false, preventing a normal CANCELED screen from replacing the emergency HMI screen.
  • EMERGENCY_STARTED and EMERGENCY_CLEARED notify robots in every state, not only active robots.
  • One HMI notification failure cannot block the fleet workflow; notification uses fire-and-forget failure isolation.
  • notifiedRobots contains robot names for which notification was attempted. It does not mean every HMI acknowledged delivery.
  • An emergency return may bypass an unbacked stale reservation, but not tracking that clearly belongs to another task.
  • Mission continuation first captures the current emergency versions and clears only matching return requests and fleet pauses. A newer version created concurrently must remain active, and EMERGENCY_CLEARED must not be broadcast.

7.1 Emergency Version Concurrency Rules

  • Every ReturnToErPharmacy event creates a new version. Robots that are already paused keep their original taskAcceptancePausedAt, while their version advances to the newest event.
  • MissionContinues captures all current versions before conditional clearing. An emergency that begins after the capture does not match and cannot be cleared by the older request.
  • Older rows without a version are treated as the null legacy version so the first clear after upgrade can remove them safely.
  • EMERGENCY_CLEARED is published and broadcast only after confirming that no paused robots remain.

7.2 Emergency Return Idempotency and Recovery

The MongoDB EmergencyReturns collection uses the original taskHistoryId as its unique _id. Concurrent HMI continues or redelivered events for one task have exactly one operation owner. A replay after completion returns the same returnTaskId.

statephaseMeaning and action
PROCESSINGCLAIMEDCancellation intent has not been recorded; inspect backend logs and Mongo state if it remains here
PROCESSING / FAILEDCANCELLATION_REQUESTEDRMF cancellation may not have been sent, may have been rejected, or may be unknown; reconcile with RMF before any replay
PROCESSING / FAILEDORIGINAL_CANCELEDA legacy or not-yet-allocated flow; it cannot prove that no return was created before a crash, so only manual reconciliation is allowed
PROCESSING / FAILEDRETURN_ALLOCATEDreturnTaskId was persisted first; TaskHistory and TaskQueue can be completed idempotently with that same ID
COMPLETEDRETURN_QUEUEDThe return exists; use returnTaskId for subsequent tracking

The API never automatically reclaims a FAILED operation. On-call staff must reconcile the RMF booking, original TaskHistory, robot ownership, return TaskHistory, and TaskQueue before following the site procedure to create a return manually or correct the operation. Deleting an EmergencyReturns record re-enables the side effect and is forbidden until reconciliation is complete.

Run yarn diagnose:emergency-returns to list operations that have remained unresolved for more than five minutes and view the operation, original task, known return task, queues, and robot ownership together. Use --stale-minutes 0 for every unresolved operation or --task-id <original-task-history-id> for one exact operation. This command is MongoDB read-only and always reports automaticReplayAllowed: false; recommendedAction classifies the reconciliation work and never authorizes an automatic recovery.

The new flow persists a preallocated returnTaskId in RETURN_ALLOCATED before it creates return TaskHistory. Only FAILED + RETURN_ALLOCATED is eligible for the guarded recovery command. The command uses that same ID to fill in a missing TaskHistory or TaskQueue and never replays RMF cancellation. Copy the exact updatedAt from a fresh diagnostic and provide the operator, reason, and confirmation flag. A conditional Mongo update prevents a stale diagnostic or concurrent operator from obtaining recovery ownership, and every result is appended to recoveryAudit.

8. Data Authority and Consistency

DataAuthoritative sourceNotes
Task history and task ownershipMongoDBRedis or RMF events do not replace task history
Queue lifecycle, leases, and attemptsMongoDBClaim and update operations must remain atomic
taskAcceptancePausedAt and taskAcceptancePauseVersionMongoDBFleet-wide emergency gate and conditional clear
returnToErPharmacyRequestedAt and emergencyRequestVersionMongoDBTask-level emergency request waiting for HMI continue
Emergency return operationMongoDB EmergencyReturnsOriginal-task idempotency, cancellation phase, and return task id
RMF status, battery, and locationRedis runtime stateHas a TTL and may be missing
fullCharging and obstacle telemetryRedis runtime stateExposed as an overlay by GET robots
HMI busy leaseRedisMust be included in availability decisions
RMF booking reconciliationMongo task/queue plus RMF inboundRe-delivered inbound events must be idempotent

Redis reads must distinguish between:

  • missing: no runtime state currently exists for the robot.
  • failure: Redis could not be read.

Dispatch, availability, auto-charge, and emergency task evaluation must fail closed on failure. A read failure must never be interpreted as an idle, available, or empty state. A Mongo repository failure must not be mapped to “robot not found.”

9. Queue and Dispatch Safety Rules

The queue lifecycle is:

PENDING → PROCESSING → DISPATCHING → IN_PROGRESS → COMPLETED / FAILED

All contributors must preserve these rules:

  • A worker may process only the record it successfully claimed through an atomic MongoDB findOneAndUpdate.
  • Claiming writes workerId, leaseExpiresAt, and attemptCount.
  • DISPATCHING must be persisted before calling RMF.
  • An explicit RMF rejection may be retried within the configured attempt limit.
  • An ambiguous network result must not be dispatched automatically again, because RMF may already have created the task.
  • Expired PROCESSING leases may be safely reclaimed; unreconciled DISPATCHING records move to an explicit failed state.
  • taskHistoryId is the idempotency root.
  • A multi-station task reuses the same queue record and increments dispatchSequence.
  • Robot Keeper must create charging tasks with an atomic deduplication condition.
  • An ordinary task without docking map metadata must still be dispatchable. Charging metadata is required only when the task genuinely needs docking activity.

10. Dependency Injection and Port Binding

Inject a port through its Symbol token, not through a concrete adapter class:

constructor(
  @Inject(FLEET_COMMAND_PORT)
  private readonly fleetCommand: FleetCommandPort,
) {}

Bind the adapter with useExisting in the adapter module:

RmfFleetAdapter,
{
  provide: FLEET_COMMAND_PORT,
  useExisting: RmfFleetAdapter,
}

useExisting ensures that the token and concrete provider resolve to the same singleton instead of creating the adapter twice.

When adding a port:

  1. Place it in application/ports of the feature that owns the requirement.
  2. Define transport-neutral commands and results.
  3. Implement it in an adapter.
  4. Bind it with useExisting in AdaptersModule and export the token.
  5. Test the application service with a fake port or mock.

Place a contract in operations/ports only when it is shared by different top-level modules and represents explicit cross-module collaboration.

11. Common Change Scenarios

11.1 Adding a General Task Rule

  1. Put the invariant in operations/task/domain or the task application use case.
  2. If robot or fleet data is required, depend on an existing contract; do not import a Mongo repository class.
  3. If RMF needs a new capability, extend FleetCommandPort and the RMF adapter.
  4. Add task service tests and, when dispatch is involved, queue and RMF adapter tests.

11.2 Adding a Hospital Event

  1. Validate the transport request in a Hospital DTO.
  2. Interpret the hospital event in csh-hospital/workflows.
  3. Use public Operations capabilities for task and robot actions.
  4. If Operations lacks a complete use case, add one instead of letting Hospital modify more repository fields directly.
  5. Preserve HTTP status codes, error codes, and response shapes unless an API version change has been formally agreed.

11.3 Adding Another Hospital

Do not copy Operations. Create another Hospital module, implement its site configuration and orchestration workflow, and reuse EmergencyOperationsPort plus feature public capabilities.

Before integrating another hospital, identify:

  • Station naming and return-station rules.
  • Emergency broadcast and HMI endpoint differences.
  • Authentication and HIS/CMP payloads.
  • Medication and employee data sources.
  • Whether medicine-cart continue needs a different Hospital policy or workflow.
  • How AppModule selects exactly one Hospital module and site configuration for a single-site deployment.

11.4 Adding or Replacing an External Integration

  • Keep Axios, ROS, and Socket.IO clients in adapters.
  • Inject typed configuration into adapters; do not read raw environment variables in application code.
  • Map external failures to results or errors understood by the application layer.
  • Normalize raw RMF payloads in the inbound adapter before calling UpdateRobotStateService.

11.5 Changing an HTTP DTO or Route

  • Keep DTOs in the API layer. Do not add Swagger or class-validator decorators to domain models.
  • Keep application commands/results separate from HTTP DTOs.
  • Preserve existing routes, status codes, error codes, and request/response shapes. If a breaking change is unavoidable, introduce a new API version and coordinate with consumer teams.
  • Compare Swagger paths, methods, statuses, and schemas after the change.

12. Import Rules

Allowed

// A controller uses an application capability through the feature public entry.
import { TaskRequestService } from '../operations/task/public';

// An application service injects its own or another feature's contract.
import {
  ROBOT_STATE_STORE,
  RobotStateStore,
} from '../../robot/application/ports/robot-state-store.port';

Forbidden

// A controller bypasses the public entry and imports the feature internals.
import { TaskRequestService } from '../operations/task/application/task-request.service';

// An application service imports a concrete adapter.
import { RmfFleetAdapter } from '../../../adapters/rmf/rmf-fleet.adapter';

// Domain code depends on NestJS, Mongoose, or Swagger.
import { Injectable } from '@nestjs/common';

The current .eslintrc.js checks that:

  • Domain code does not import frameworks or infrastructure.
  • Application services, ports, and workflows do not import adapters, infrastructure, Product API, Hospital implementations, or concrete transport packages such as Axios, Mongoose, Redis, ROSLIB, or Socket.IO.
  • Controllers do not import repositories, adapters, infrastructure, or feature-internal application trees.

test/service/import-boundaries.service.spec.ts uses intentionally invalid imports to verify the ESLint overrides. Run it whenever override order or globs change so a later Hospital override cannot silently replace controller restrictions.

Architecture still requires reviewer judgment. ESLint can check import patterns, but it cannot determine whether a service exposes too much internal capability.

13. Current Boundaries and Further Convergence

13.1 The Emergency Facade Is the Single Cross-Module Entry

HospitalMissionControlService no longer injects feature-local repository or control ports. It uses only behavior-oriented capabilities from EmergencyOperationsPort:

  • listActiveTaskContexts
  • broadcastEmergency
  • pauseFleetTaskAcceptance / clearFleetTaskAcceptance
  • markReturnRequested / clearReturnRequested / markReturnHandled
  • cancelAndQueueReturn

EmergencyOperationsService implements this port inside Operations and coordinates task, robot, fleet, runtime state, HMI, and queue behavior. When adding emergency capability, expose a complete business operation instead of repository CRUD.

13.2 public/index.ts Is an Enforced Cross-Module Import Boundary

Hospital may use explicitly exported capabilities such as TaskProgressionService from operations/task/public, but it must not import operations/*/application/**. .eslintrc.js enforces this rule for src/csh-hospital/**/*.ts.

13.3 Task Progression and Hospital Orchestration Have Separate Registration

TaskProgressionService is a general task application capability registered by TaskModule. MedicineCartTaskWorkflow is CSH Hospital orchestration registered and exported by CshHospitalModule for Product API controllers.

Do not move Hospital policy back into TaskProgressionService, and do not make Operations depend on CshHospitalModule. New branch outcomes should remain transport-neutral and be interpreted by the Hospital workflow.

13.4 Each Deployment Currently Uses One Hospital Policy Set

The current composition expects each deployment to select one Hospital module and one SiteOperationsConfig. Supporting several hospitals in one process requires site/hospital routing; one global configuration token is insufficient.

14. Testing and Acceptance

14.1 Minimum Validation Matrix

Change scopeRequired validation
Pure domain/value objectLint, build, feature unit tests
Application serviceAbove plus service tests
Module/provider/public entryAbove plus module-boundary tests
Port/adapterAbove plus adapter contract tests
Queue/emergency/runtime stateMongoDB/Redis integration tests plus E2E
HTTP DTO/routeE2E plus Swagger contract comparison
Deployment/configurationProduction image build plus startup smoke test

14.2 Standard Commands

Use Node.js 24.13.0 to match the pipeline:

yarn install --frozen-lockfile
yarn lint:check
yarn test:architecture
yarn build
yarn test test/service --runInBand
yarn test:integration --runInBand
yarn test:e2e --runInBand

Service integration tests and E2E require MongoDB, Redis, and a test .env. Connection failures in a bare checkout without these dependencies are neither program regressions nor successful validation.

The Jest configurations intentionally do not use forceExit. A process that prints test results but does not terminate is still a lifecycle failure. E2E suites must close the Nest application through AppHelper.closeAgent(), test databases through MongoHelper.close() so the underlying pool is closed, and new long-running adapters or workers must abort or drain work in a Nest shutdown hook.

Docker Compose validation is recommended:

docker compose up -d mongodb redis
docker compose run --rm backend yarn test test/service --runInBand
docker compose run --rm backend yarn test:e2e --runInBand

The effective Compose configuration must point the test container to the mongodb and redis service names, not to localhost inside the container.

14.3 Critical Regression Cases

Any Operations/Hospital change should verify at least:

  • An ordinary single-station task without docking metadata can still be dispatched.
  • Only one worker wins a multi-worker claim race.
  • Dispatch timeout or an unknown result does not trigger redispatch.
  • Multi-station continuation increments the sequence correctly.
  • Duplicate RMF terminal events do not complete a task twice.
  • Redis failure causes dispatch and auto-charge to fail closed.
  • Redis or Mongo read failure makes emergency evaluation fail closed while the fleet pause remains active.
  • Emergency pause blocks normal tasks and Robot Keeper.
  • An older MissionContinues cannot clear a newer emergency version.
  • Concurrent emergency returns for one original task produce only one cancel/create side effect.
  • RMF cancel rejection or timeout does not delete the local queue or create a return; the operation phase remains available for reconciliation.
  • Emergency return follows the stale-reservation rules.
  • Manual cancellation remains rejected during emergency pause.
  • Fleet-wide HMI broadcast includes robots in every state and preserves notifiedRobots.
  • Obstacle and full-charging telemetry overlays do not regress.
  • Docking-arrival decision-manager activity and activity order remain unchanged.

15. Pull Request Collaboration Process

15.1 Author Checklist

  • The owning feature or Hospital boundary has been identified.
  • No hospital name or hospital-specific environment condition was added to Operations.
  • Controllers perform only validation, mapping, application calls, and response mapping.
  • Application code does not import a concrete adapter.
  • New ports use string IDs and transport-neutral types.
  • Raw RMF, Axios, Mongoose, or Redis representations do not cross adapter boundaries.
  • Queue, emergency, idempotency, and fail-closed invariants remain intact.
  • Relevant unit, integration, and E2E tests pass.
  • Documentation and .env.example are updated for API or configuration changes.

15.2 Reviewer Checklist

  • Business rules are placed by ownership, not by the source of the request.
  • Any newly public service genuinely needs to be exported from the feature’s public entry.
  • Hospital has not gained another repository-port dependency; if it has, consider an Operations use case instead.
  • Operations has not gained a dependency on a concrete CSH/CMP/HIS implementation.
  • Error mapping and HTTP contracts remain compatible with existing consumers.
  • Failure and ambiguous outcomes are distinguished correctly.
  • Fire-and-forget is used only for notifications that permit partial failure.
  • Data authority and transaction/atomic-update behavior are explicit.

15.3 Merge and Deployment

  • Each commit should build and be independently revertible.
  • Keep data migrations separate from code-only refactors.
  • Run queue schema migrations according to task-queue-v2-migration-runbook.md: stop the old worker, migrate, and only then start the new version.
  • After starting the production image, follow deployment-runbook.md to validate MongoDB, Redis, RMF, Fleet Manager, HMI, and Hospital integrations.

16. Suggested Team Ownership

AreaPrimary ownerRequired collaborators
Operations domain/applicationBackend/Operations teamRMF, Hospital, QA
CSH Hospital workflows/integrationsHospital Integration teamBackend, Hospital API owner, QA
RMF/Robot adaptersRobot Integration teamBackend, fleet vendor
Product/Hospital HTTP contractAPI ownerConsumer teams, QA
MongoDB/Redis schemas and migrationsBackend/Data ownerSRE, QA
Deployment/configurationSRE/PlatformBackend, Hospital IT

Before starting cross-team work, identify at least the owner, API/port contract, data authority, failure semantics, idempotency strategy, test environment, and deployment order.

17. Important File Index

FilePurpose
src/operations/operations.module.tsOperations feature composition
src/operations/*/*.module.tsFeature providers and exports
src/operations/*/public/index.tsFeature public import entries
src/operations/ports/Cross-top-level-module contracts
src/operations/workflows/Emergency facade, Queue Worker, and Robot Keeper
src/csh-hospital/csh-hospital.module.tsHospital providers and workflow composition
src/csh-hospital/workflows/hospital-mission-control.service.tsEmergency and return-to-pharmacy workflow
src/csh-hospital/workflows/medicine-cart-task.workflow.tsHMI task progression and emergency-return orchestration
src/adapters/adapters.module.tsComposition of ports and concrete adapters
.eslintrc.jsStatic import-boundary rules
test/service/module-boundaries.service.spec.tsModule-structure and provider-graph validation
docs/modular-monolith-architecture.mdSystem architecture summary
docs/deployment-runbook.mdBuild, deployment, and rollback instructions
docs/task-queue-v2-migration-runbook.mdQueue migration procedure

18. Glossary

  • Feature public entry: A feature’s public/index.ts, which explicitly lists externally usable application services.
  • Port: An abstract capability required by application code and injected through a NestJS token.
  • Adapter: A technical implementation of a port, such as a MongoDB repository or RMF client.
  • Workflow: A business process coordinating several use cases or features.
  • Runtime state: Current RMF/robot state stored primarily in Redis.
  • Task ownership: The task currently assigned to a robot, with MongoDB as the authority.
  • Fail closed: Reject dispatch or automation when external state cannot be verified instead of assuming it is safe.
  • Dispatch reconciliation: Determining the actual dispatch result from inbound RMF booking/status data after an ambiguous response.
  • Idempotency root: The stable identifier used to recognize one business task; currently taskHistoryId.