---
title: Engineering for Resiliency - Ximedes
description: The website of Ximedes
---

[Ximedes Blog ](https://ximedes.com/blog)

# [Engineering for Resiliency - Ximedes](https://ximedes.com/blog/2021-03-15/sim-basics)

 Written by [Rick van Krevelen](https://ximedes.com/blog/author/rick-van-krevelen) | 14/03/2021

Money transfers, fare payments, fleet movements: **Ximedes** plays a central role in realizing core processes where [resiliency](https://www.wikiwand.com/en/Resilience#/Technology_and_engineering) is critical. Achieving system resiliency, that is, ensuring adequate quality of service while being fault tolerant, quickly becomes tricky due to the inherent complexity of such cyber-physical systems. With critical business processes operating across disparate platforms and devices, how would one go about engineering resiliency to emerge "naturally", similar to *herd immunity* and other emergent phenomena?

**TL;DR** — Extending system test scenarios with simulation basics (common clock and reproducible randomness) can help to rigorously analyse the complex contexts that trigger undesired system behaviors to emerge, and help to prevent them.

## Hold the Chimps

Although important, "standard" functional testing may fall short in verifying a system's resiliency to system-level risks, particularly risks regarding the physical portion of the cyber-physical systems mentioned earlier. To maintain sanity and prevent regression, business logic typically undergoes verification & validation continuously, and perhaps also the occasional (costly) accreditation during development.

A myriad of software testing methods will help to achieve satisfactory levels of code coverage and other quality assurance metrics across many use cases beyond merely the happy flow. Unfortunately, consistently reproducing any (non-functional and/or physical) disturbances and disruptions from which the system may need to recover [autonomically](https://www.wikiwand.com/en/Autonomic_computing), could take testers a bit more work.

Companies like Netflix for instance apply [chaos engineering](https://www.wikiwand.com/en/Chaos_engineering), unleashing a full *Simian Army* to wreak havoc on production environments and so help improve system resiliency before final release. Once released, remaining undesired behaviors may be mitigated or even resolved during quality control and after-sales service. Presumably however, one is better safe than sorry.

At Ximedes, we understand the challenges involved in handling critical events in large-scale platforms. Examples include our omni-channel PSP services for merchants ([OmniKassa](https://ximedes.com/2018-04-12/rabobank-omnikassa/)), controlling a fleet of public transport vehicles ([GIVA](https://ximedes.com/2018-02-23/pto-control-vehicle-operations/)), and seamless payments for [vending machines](https://www.pecunda.com), [fares](https://www.tapconnect.io), [unattended shops](https://ximedes.com/2019-09-05/ximedes-powers-ing-seamless-payments-for-ah/), etc. Supporting the primary business process, these mission-critical systems require not only [secure coding](https://ximedes.com/2018-04-07/secure-coding-at-ximedes/) throughout (enroll your team for a course [here](https://securecoding.ximedes.com/)). To validate their robustness, extended testing practices are also required.

## Determinism

In order to analyse system behaviors with rigor, we would like to reproduce the exact steps that culminated in notable (undesired) emergent phenomena in the system under test. Engineering deterministic system behavior is however not straightforward: concurrent events occurring across cloud-based and mobile devices are preferably treated [asynchronously](https://www.wikiwand.com/en/Asynchronous_system) in a [reactive](https://www.wikiwand.com/en/Reactive_programming) manner while avoiding the need to synchronize behaviors (across threads, actors, nodes, etc.) to prevent [deadlocks](https://www.wikiwand.com/en/Deadlock).

In cases where synchronization across disparate system components is critical, coordination languages help to effectively re-introduce determinism by making explicit the timing assumptions along with their associated (fault) behaviors. [Lingua Franca](https://github.com/icyphy/lingua-franca) for instance extends the *C* and *TypeScript* programming languages (others to follow) with keywords like `delay` and `deadline`, applying universal coordination principles based on simulation formalisms like [DEVS](https://www.wikiwand.com/en/DEVS) and standards like [HLA](https://www.wikiwand.com/en/High_Level_Architecture).

Even without time coordination, just using *events* or an event-driven architecture ([EDA](https://www.wikiwand.com/en/Event-driven_architecture)) for the business logic can prove helpful, particularly when history and scale are essential. By guaranteeing that all changes to domain-specific objects are initiated by event objects, the system not only

1. provides an **event log** suitable for auditing and debugging the intent or reason of state changes, but also
2. allows **temporal querying** of the system's state across multiple, altered timelines, and even
3. enables **system recovery** via event replay, possibly starting from recent snapshots.

The [event sourcing](https://martinfowler.com/eaaDev/EventSourcing.html) approach for instance achieves this by organizing all domain event handling into *aggregates* and *views*. Since retrofitting such coordination approaches into your business logic is often cumbersome, difficult, or simply impossible, one had best apply these approaches early on.

There is however a good chance your system is already mature and shows nondeterministic behavior, for instance varying execution and communication performance due to mobile or cloud deployment with asynchronous event handling. Fortunately, you can still benefit from determinism: in the system's simulated test environment!

## Modeling discretely

Model-driven engineering ([MDE](https://www.wikiwand.com/en/Model-driven_engineering)) processes typically involve delivering a set of system test scenarios. Perhaps these could be suited beyond just verifying the business logic integrity, and also help stress-testing its resiliency, by applying some simulation techniques. But what techniques or simulation type should we apply?

**Continuous simulation** models consist of ordinary differential equations (ODEs), but lack stochastic transitions as well as discrete (integer) counts. Examples are *stock-and-flow* models such as the [beer distribution game](https://www.wikiwand.com/en/Beer_distribution_game) for demonstrating the [bullwhip effect](https://www.wikiwand.com/en/Bullwhip_effect) occurring in supply chains, and *compartmental models* that approximate susceptible-infectious-recovered ([SIR](https://www.wikiwand.com/en/Compartmental_models_in_epidemiology)) proportions of a population during an epidemic.

Conversely, **discrete-time simulations** lack continuous time and are stuck in fixed-length cycles or ticks, as for example in *turn-based games* and *cellular automata* such as [Conway's Game of Life](https://www.wikiwand.com/en/Conway%27s_Game_of_Life) and simpler ticker tape-like automata.

Alternatively, **discrete-event simulation** ([DES](https://www.wikiwand.com/en/Discrete-event_simulation)) is generally a good choice, as this approach yields reproducible simulation models that have it all: continuous time, discrete counts *and* stochastic state transitions. The main ingredients for a reproducible or deterministic DES-type simulation scenario are:

- a *common* clock or [time reference](https://www.wikiwand.com/en/Frame_of_reference) which is advanced by a [priority queue](https://www.wikiwand.com/en/Priority_queue) of time-ordered test events;
- a *common* [pseudo-random number generator](https://www.wikiwand.com/en/Pseudorandom_number_generator) for all [probability distributions](https://www.wikiwand.com/en/Probability_distribution) representing uncertainty in your test model.

## Stay tuned

In a follow-up we will explore *how* these simulation techniques could be applied with little effort into existing system integration tests, to help engineer some reproducible chaos, and fix undesired behaviors *before* deploying to production.

Want to know more about how **Ximedes** can help you evaluate the resiliency of your system? [Contact us](https://ximedes.com/contact)!

[View full post](https://ximedes.com/blog/2021-03-15/sim-basics)

```json
{
  "@context" : "http://schema.org",
  "@type" : "BlogPosting",
  "author" : {
    "@type" : "Person",
    "name" : "Rick van Krevelen"
  },
  "dateModified" : "2025-10-17T09:25:00.394Z",
  "datePublished" : "2021-03-14T23:00:00Z",
  "headline" : "Engineering for Resiliency",
  "image" : {
    "@type" : "ImageObject",
    "height" : 480,
    "url" : "https://2212407.fs1.hubspotusercontent-na1.net/hubfs/2212407/Blog%20Images/persistence_of_memory.jpg",
    "width" : 720
  },
  "mainEntityOfPage" : "https://ximedes.com/blog/2021-03-15/sim-basics",
  "publisher" : {
    "@type" : "Organization",
    "logo" : {
      "@type" : "ImageObject",
      "height" : 60,
      "url" : "/hs/hsstatic/content_shared_assets/static-1.4092/img/default-amp-logo.png",
      "width" : 60
    },
    "name" : "Ximedes Blog"
  }
}
```