🔒 Private · Internal

ITSCM Learning Hub

This space is private. Enter the passphrase to continue.

Tip: this gate only hides the page from casual view. For real protection, also switch on password protection at the hosting level (see the note your colleague left in the footer).
ITSCM Learning Hub
Knowledge base · built from the handover sessions

Everything you need to run ITSCM testing at Swiss Life AM.

A calm, structured place to learn the role before your predecessor leaves — the concepts, the day-to-day workflow, and exactly how the tools (Jira, Xray, Confluence) fit together. Work through it in order, or jump to what you need.

2 strategy documents 3 handover sessions 9 diagrams 5 checklists 90-day playbook 40-day study plan
00 · Start here

Your job, in one sentence

You make sure Swiss Life AM can recover its critical IT services after a disaster — and you prove it works by testing, so the company satisfies regulators (DORA, FINMA) and avoids a multi-million recovery bill. You don't run every test yourself; you organise, prepare, and document, and you support the people who own the systems.

50 / 50
Your split: ITSCM testing + supporting Roland
5
Building blocks: Epic · Story · Test Plan · Test Cases · Report
1 / 3 yr
Test cadence: yearly, or every 3 years for DORA
The 6 things to know cold

1. RTO = max downtime · RPO = max data loss  ·  2. S1 = no data loss · S2 = data loss · S3 = ransomware  ·  3. O1 fixes S1 · O2 fixes S2 · O3 fixes S3  ·  4. IRRP = recover one system · FRCP = recover everything, in order  ·  5. BCM is the boss; ITSCM is its IT arm  ·  6. DORA + FINMA require you to test recovery and prove it.

How to use this hub

  • Read sections 00–03 first for the big picture, then 04–06 for the tools and the plan you'll live in daily.
  • Tick off the checklists as you go — progress saves in your browser and shows in the sidebar.
  • Preparing before you start? Go to the 40-day study plan — one 30-minute session a day, with a calendar file for daily reminders.
  • Anything marked confirm is my best reading of a noisy recording — verify it with your colleague before relying on it.
01 · Your role

The Test Manager as an enabler

This was the core message of the handover: you are not expected to personally execute every test. You coordinate, prepare, simplify, and follow up so the application owners (the specialists) can get their systems tested without drowning in work.

Your Role: the Test Manager as an Enabler You don't run every test yourself — you make it easy for the people who own the systems to get tested. ITSCM Test Manager (YOU) — coordinate, prepare, simplify, follow up, report Application Owners (a.k.a. technical contact / specialist) they keep their expertise & steer their app support deliver The 6 working principles she stressed 1 · Support, don't demand The owners are overloaded. If you demand everything from them, nothing happens on time. Take work off them; let them keep their specialty. 2 · Learn each app deeply Understand its architecture & how its backup works. Look inside often — it's like a Russian doll: open one layer, find more. That tells you what is worth testing. 3 · Mine Jira for scenarios If a scenario is hard to define, pull an extract of the app's incidents/tickets. Real incidents show you what actually breaks — and become test input. 4 · Be persistent “Steter Tropfen höhlt den Stein” — the constant drip wears the stone. Keep coming back: “Done? How can I help?” Most people thank you for it. 5 · Use DORA as the carrot DORA is the legal lever AND a budget argument. “We must, by law” unlocks time, money and cooperation you couldn't get otherwise. 6 · Frame it as win–win Scenarios that work = owners look good, management is reassured, the regulator is satisfied, and the company avoids a multi-million recovery bill. Background that helps in this role • Infrastructure & network thinking — you need to “think in systems and dependencies,” not just one app. • Knowing the desired end-result (what “recovered” looks like) makes the whole support chain easier to steer. • Adaptability — you adapt to each specialist (even down to speaking their language: High German vs. Bernese). She described her old job as “playing the firefighter”; this role is the opposite — prepare in advance so the fire never spreads.
Diagram · Your role and the six working principles
Her teaching philosophy

"Don't give them a fish — show them how to fish." You give owners a repeatable structure and a nudge, not a finished test. Most people thank you for the support; a few need patience. Persistence wins: steter Tropfen höhlt den Stein.

02 · Core concepts

RTO, RPO, scenarios & solutions

Two numbers and four scenarios sit under everything. Master these and the documents read themselves.

RTO vs RPO — The Recovery Timeline The two numbers that drive every protection class. Look at where each one sits on the timeline. Last good backup ⚡ DISASTER happens here Service restored RPO Recovery POINT Objective "How much DATA can we lose?" = the gap since the last backup RTO Recovery TIME Objective "How long can we be DOWN?" = downtime until service is back 2h crisis decl. (part of RTO) Smaller numbers = more expensive. A demanding RPO means backing up very often (lose almost nothing); a demanding RTO means recovering very fast. SLAM's protection classes set both numbers per application.
Diagram · RTO measures forward (time to recover); RPO measures backward (data lost)
ITSCM: Outage Scenarios → Solutions Each failure type maps to a specific technical solution. The solutions stack: O3 builds on O2. THE FAILURE (Scenario) WHAT HAPPENED THE FIX (Solution) S0 Single service/app fails One app goes down. Normal IT ops handle it. OUT OF ITSCM SCOPE S1 Infra fails — NO data loss Building/system died, but data is intact. O1 — Geo-redundant Synced systems → fail over to the 2nd data center S2 Infra fails — WITH data loss The data itself was lost or damaged. O2 — Geo-redundant backup Restore from backup copies held at both sites S3 RANSOMWARE data encrypted by attacker Attacker encrypts data AND tries to destroy the backups too. O3 — Immutable backup WORM storage attackers can't change or delete (recovery can take days) Remember: S1 = switch sites · S2 = restore backup · S3 = restore from a backup the attacker couldn't touch. All RTO targets assume the crisis is declared during business hours — otherwise "Best Effort" applies. The RTO clock includes a 2-hour crisis-declaration phase.
Diagram · Each failure scenario maps to a solution; the solutions stack (O3 builds on O2)
O1 · fixes S1

Geo-redundant sync

Two data centres, data synced. If one dies, fail over — no data loss.

O2 · fixes S2

Geo-redundant backup

Backups at both sites. Restore after data loss. Mandatory for every productive app.

O3 · fixes S3

Immutable backup

WORM storage attackers can't change or delete. The ransomware insurance.

03 · How a test works

From "we should test this" to a green tick

The practical workflow she walked through on screen. The golden rule: get as close as possible to a real disaster scenario without touching production — you almost always build and run in a UAT / test environment.

How a Real ITSCM Test Gets Created (in practice) The day-to-day workflow she walked through — the part that is not written in the formal documents. 1 Pick a business-critical app From the BIA / inventory — criticality decides whether it needs an ITSCM test at all. 2 Understand how it works Its architecture AND its backup. “Look inside often” — it's a Russian doll; each layer reveals what is worth testing. Do this WITH the application owner. 3 Define the scenario (S1 / S2 / S3) Agree what you simulate: e.g. S1 = the site/app no longer runs. If the scenario is hard to define → pull a Jira extract of the app's incidents to see what really breaks. 4 Build a test environment & run it Stand it up in a test/UAT environment, execute the restart or recovery, and watch how it behaves. 5 Measure against the RTO — and prove it How long did recovery take? Did it meet the target? Capture the evidence. 6 Document & report Log it in Jira / Xray (Epics, User Stories, bugs + fixes). The per-test report is today a Word document. Shortcut: document real events as tests Keep your eyes & ears open. When a real restore, backup, or DC switch happens (often over a weekend, after dev work), you can document it retroactively as an ITSCM test. Ask teams to tell you — it saves building a separate test from scratch. How often (cadence) Yearly for restart/recovery plans (per AM 12.1.2). Every 3 years is the DORA rhythm — but the world changes, so re-check what's still valid. On a significant change (e.g. new backup infra), run the test at the change — sometimes only a theoretical test if a practical one isn't feasible.
Diagram · How a real ITSCM test gets created in practice

Two worked examples she showed

Bloomberg · practical

Connection test in UAT

Precondition: network & security up. After the tunnel is back, the connection re-establishes. Test cases were written as steps 1-2-3-4; production was never touched. Result: successful, with one config gap logged as a defect.

Data-centre · yearly

Start vs. Recover

Two data centres — St. Gallen (primary) and Chur. A Start just brings systems back up (S1); a Recover restores the whole environment from backup (S2). Run yearly with the provider (Inventx).

Document real events as tests

Keep your ears open. When a real restore or data-centre switch happens (often over a weekend after dev work), you can document it retroactively as an ITSCM test — it saves building one from scratch. Ask teams to tell you when they happen.

04 · Jira & Xray

The tool logic: how a test is structured

Every test lives as a small tree of linked Jira items. Learn this hierarchy and you can read, build, or pick up any test. These are the five building blocks she called the heart of the job.

Epic
The test initiative — one app (or interface), one cycle

The container for everything about testing an application in a given year. Luxembourg and others click into the Epic to see how a test ran. If a related app shares an interface (e.g. Maxis + Salesforce), they can share one Epic.

User Story
What is being validated

Describes the scenario in business terms — including which of S1 / S2 / S3 applies. Reusable next cycle; you don't reinvent it each year.

Test Plan (Xray)
The container for the test cases

Groups the test cases for this initiative and tracks their execution.

Test Case
The actual steps

Step 1-2-3-4… with preconditions and expected results. This is the part she said could be auto-generated into the Word report later.

Task
Your own to-dos

e.g. "prepare the report," "chase the owner." How you track the coordination work that's yours, not the specialist's.

Defect / Bug
What failed, and the fix

Copied straight out of the failing test case (all the info comes with it). You assign it and set criticality (critical / medium). A test with an open defect is a failed test until it's fixed and re-run.

A test moves through a small set of statuses. The dashboard groups everything by these so you always know what needs your attention.

StatusMeaningYour action
To DoPlanned, not started.Plan it, agree a date with the owner/provider.
In ProgressThe test is running.Support, follow up, watch for defects.
ReportingTest done, waiting on the report.Chase the owner to fill & sign the report.
DoneComplete and closed.Mark green in the plan; inform the relevant countries.

confirm In the Excel plan, colours also flag test outcome: black = planned · red = failed (needs rework) · green = done. (Application-category colours are covered in section 06.)

  1. Talk to the app owner / provider and agree what and when (e.g. "we'll test in Q3").
  2. Create the Epic, then the User Story, the Test Plan and its Test Cases.
  3. Add a Task for yourself (usually: prepare the report).
  4. Copy the report template (based on the Swiss Life AG one, adapted to your environment).
  5. Build the environment in UAT, run the test, capture timing vs. the RTO.
Rhythm she recommended

Plan Q1 in Nov/Dec. September is quiet. Q1 is mostly re-tests (things that failed and need rework) plus small clarifications. Contact everyone from mid-January to mid-February to lock the year.

Defects are normal — early on a batch of tests can throw 15–20+. Each is copied from the failing test case, assigned, and given a criticality. You stay on them until fixed.

  • Many defects are script-related: the recovery script didn't run cleanly. The person who owns the script finds the root cause, adapts it (often alongside a new version), and you re-run and re-document the test until it goes green.
  • You can automate the nudge — log a work item; if nothing happens, send a mail.
  • A recovery-plan gap found in discussion with technicians (e.g. "responsibility for Bloomberg Terminal connections isn't mapped") also becomes a defect.

When the test is done, the report is filled by the person who knows the product (there's no lock on a shared Word file, so the owner — not you — edits it) and each participant adds a dated comment in the report's chapter 5.

Reto reviews, then it goes to the business manager. The report is signed off by the Application Manager, ITSCM Manager, Reto, and Business Manager.

The Dashboard in Jira rolls everything up by status (To Do / In Progress / Reporting / Done). Note: closed items get archived each year — archived Epics won't show in a normal search, so keep the master list of Epic links.

05 · Confluence & where things live

The map: where to find what

Almost everything is documented in Confluence, with links out to the plan (Excel), the process models (ADONIS), and templates (SharePoint). Here's the map she gave you.

WhereWhat's thereWhy you need it
Confluence · Operational ResilienceThe summary landing page tying it all together; the Strategy PDF and the Testing Framework instructions are linked/embedded.Your home page. confirm some content is being refreshed to current state.
ADONIS (Swiss Life)The BPM process models — roles, who does what, per country, the Emergency Organisation process.You'll need the "IT Testing" permission. Click a role to see its tasks.
Excel · Test Activity PlanThe master plan of all apps, tiers, countries, and scheduled tests (see section 06).Your day-to-day planning tool. Maintained by your team; Luxembourg reads it.
Jira DashboardLive status of every test initiative, grouped by status.Your "what needs attention today" view.
SharePoint · TestingTemplates (report, etc.). Some old PowerPoint/Word templates are retired — she now uses one.Grab the current report template here. A name on a file doesn't mean that person authored it.
Team channel (Teams)The working links, the document repository, and a Whiteboard with the exercise agenda & per-stage checklists.Guests (auditors) get access here. The whiteboard shows what each stage of the process requires.
Emergency Organisation

There's a dedicated process (Reto is building it and will walk you through it). You'll be registered in the org, and there are laminated cards for it. Each stage has a checklist and questions to prepare (e.g. for management support).

06 · The Test Activity Plan

The Excel master plan

One spreadsheet drives your year: every in-scope application, its tier, which countries use it, the infrastructure it runs on, and when it's tested. The goal is to consolidate all test plans into this one master over time.

Tiers & cadence

TierMeaningCadence
Tier 1Most critical apps.Tested most often (e.g. 2025, then every 3 years → 2028).
Tier 2Important.On the 3-year rhythm; sometimes pulled forward.
Tier 3Lower criticality.Around 2027 in the current cycle. confirm
New / yearlyA category they added; plus items required yearly.Yearly (e.g. the DC Start/Recover and the Full Recovery). confirm

Colour legend

confirm the exact colours — this is my reading of the walkthrough; verify against the live sheet.

Light green — Tier 1 application Yellow — has a tier, not yet planned Dark green — "quality-important" function defined
Black — test planned Red — test failed / needs rework Green — test done

Who uses each app (the country columns)

Columns flag which entities use an application — Switzerland, Germany (incl. KVG, a separate German company), Luxembourg, France, Norway. When a test finishes you use these to know which countries to inform. Contacts live in the notes column (e.g. the Norway manager).

How planning actually works

  • Each year's Epics are created ahead of time: she makes the next year's Epics around Dec/Jan, then contacts each owner/provider to confirm timing.
  • You can't just declare "I'll test X in Q1" — you agree a slot with each owner ("no time in Q1, but Q4 2027 works").
  • Future years hold placeholders (e.g. 2027, 2028) with a short comment on how the plan was derived — the workload looks like a "bathtub": heavy, then lighter, then heavy again.
Watch this

The Excel is static — after you change anything in Jira, tidy the Excel by hand so they stay in sync. A big future project (SimCorp Dimension / "SCD" as a SaaS service, release ~April) will reshape testing; Roland leads test management there. Expect the framework to slim down.

07 · The annual cycle

The November–December cycle — your hardest deadline

This is the most time-critical thing Prisca handed over. You arrive in September; the year’s biggest deadline-driven cycle begins six weeks later and runs on fixed dates.

Her clearest instruction

If you start in September, get into this by early October at the latest — better earlier. Scoping, verifying the Control Procedures and booking the walkthrough dates all have to happen before 1 November.

The risk if you don’t: on 1 November a whole battery of Control Procedures fires automatically at ~30 Application Owners who, in her words, “definitely won’t know what it’s about any more.”

The Annual ITSCM Cycle — and where you land You start in September. The busiest, most deadline-driven part of the year begins six weeks later. YOUR CRITICAL PATH Oct → early Dec: everything must be prepared and run Jan–Mar Apr–Jun Jul–Aug Sep Dec–Jan SEPTEMBER — you arrive • Get access: Jira, Confluence, SharePoint, Fact24, ADONIS • Get put on the Two-Pager as Management Support • Replace Prisca as Control Owner/Performer in Jira EARLY OCTOBER — prepare • Scoping: which apps are CIF? • Check every Control Procedure is correct before it fires • Book the walkthrough dates in people's calendars — early! • Warn owners it's coming 1–30 NOVEMBER — review 1 Nov: Control Procedures fire to ~30 Application Owners • They review their IRRP (x.9) • You + Reto cross-check & approve • Final version becomes x.0 30 Nov: all must be done WEEK 1 DEC — walkthrough • Full Recovery Plan walkthrough • Check sequence & dependencies • Only CIF = "Yes" apps reviewed • Produces FRCP v3.0 + fixes • Then upload everything to Fact24 (manual!) The rest of the year Jan–Feb: contact all owners and lock the year's test plan. Q1 is mostly re-tests of what failed. All year: run the scheduled ITSCM tests — and create new IRRPs as soon as a new CIF app appears (e.g. Axioma), or when one IRRP needs splitting (e.g. Bloomberg → Terminal + FIX Connections). Never leave this to November. Why early preparation is non-negotiable If nothing is prepared by 1 November, a whole battery of Control Procedures fires at ~30 people who will have no idea what to do. Prisca's words: get in by early October at the latest — better earlier. November is only for checking what already exists — did links, people, or support contacts change? Anything genuinely new must be built during the year, not squeezed into the review window.
Diagram · The annual cycle and where you land in it
WhenWhat happensYour part
Early OctPreparation window.Scope the CIF list, verify every Control Procedure, book the walkthrough dates, warn the owners.
1 NovControl Procedures fire automatically to ~30 owners (+ deputies).Confirm they actually went out; send a short explanatory note.
1–30 NovOwners review their IRRP. They have 4 weeks.Cross-check each one, get Reto’s approval, upload the PDF to Fact24, chase stragglers.
Week 1 DecThe Full Recovery Plan walkthrough.Run it — in four groups of ~1.5h, not two of 3h.

Last year most Control Procedures went out on 4 Nov; this year it will be 1 Nov. Two were piloted first.

The trap

Never let new work land in November. November is only for checking what already exists — did links, people or support contacts change? Anything genuinely new must be built during the year: Bloomberg had to be split into Terminal + FIX Connections, and Axioma became a new CIF with no IRRP at all. Both were handled early, so November was just a glance.

08 · IRRP & Control Procedures

How the review actually flows

Every CIF (Critical Important Function) — about 30 applications, all listed in the Excel plan — has its IRRP reviewed once a year. The mechanism that drives it is a Control Procedure in Jira.

The IRRP Review — how a Control Procedure flows Runs every November for each Critical Important Function (~30 apps). This is your core recurring duty. JIRA (Control Proc.) APPLICATION OWNER YOU (+ Reto) FACT24 (fallback store) 1 1 Nov: fires Auto-sent to the owner + their deputy. 4 weeks. 2 Reviews the IRRP In track-changes mode: • contacts & responsibilities • Restart + Recovery schedule • all links still work Saves as version x.9 3 Send to Review Clicks the button in the Control Procedure. Also supplies a PDF. 4 Your cross-check Read "diagonally": what changed, is it plausible, does it match BIA/RTO/RPO? NOT your job: server names, technical detail. Track-changes shows the diff. 5 Reto approves Four-eyes principle. Accept changes → final version x.0 Then close the Control Proc. 6 Upload the PDF Into Fact24 (SIM module). MANUAL — there is no sync. 7 First week of December — the Full Recovery walkthrough Go back into the IRRPs together, confirm the recovery sequence and dependencies, and check the RPO/RTO can actually be met. Produces the next FRCP version. The version convention x.9 = owner has reviewed, awaiting approval. x.0 = approved and final. e.g. 1.9 → reviewed → approved → 2.0. SharePoint versioning is deliberately off. Who owns what The IRRP belongs to the Application Owner. You manage that it gets reviewed, questioned, approved and stored. Every Control Procedure needs two named people (owner + deputy).
Diagram · The Control Procedure → IRRP review flow

A Jira mechanism built by Romano (Atlassian team) — not just for ITSCM. Many people use it for recurring obligations: it reminds them, they complete it, and it becomes audit evidence.

  • Starts 1 November, deadline 30 November — four weeks.
  • Contains a generic link to the IRRP folder (generic so each procedure doesn’t need editing).
  • Every Control Procedure needs two named people — a Control Performer and a deputy — so it survives someone leaving.
  • A new CIF app means a new Control Procedure. Ask Romano — he built it, he’s proud of it, and you get full support.
  • They all hang off one internal control (ICS): IT Restart and Recovery Plan and Full Recovery Plan.

Prisca was refreshingly blunt about the limits of your role.

Yours

Read diagonally

What changed? Does it exist? Is it plausible? Does it match the BIA / RTO / RPO? Do the links still work? Track-changes shows you the diff against last year.

Not yours

Technical depth

Server names and configuration detail are the owner’s responsibility. In her words: “whether the server is called that — I don’t care. That’s not my responsibility.”

The IRRP belongs to the Application Owner. You manage that it gets reviewed, questioned, approved and stored. Reto, as ITSCM Manager, digs much deeper than you will.

What the owner must check: IT service information and contacts, responsibilities, the Restart Schedule, the Recovery Schedule (3.2), that links work, and the return-to-normal-operations steps.

VersionMeaning
x.9The owner has reviewed it. Status Reviewed/Changed, with their comment and name on page two. Awaiting approval.
x.0Approved by Reto (four-eyes principle). Final. e.g. 1.9 → reviewed → approved → 2.0.

SharePoint versioning is deliberately switched off — they consciously moved away from keeping versions. You see what changed through track-changes instead.

Don’t forget the last step

Every approved PDF must be uploaded to Fact24 manually. Nothing synchronises. If it changed, re-upload it.

The Full Recovery Plan is at v2.0, produced from last December’s walkthrough. Red entries are open items for the next one — there are always some. Next time: copy it to v3.0 to keep the history.

  • Group 1 — all of Shared IT together, from the IT basis infrastructure (Recovery Class 10) upward. Keep them in one room: their cross-talk is where the value is (“no, if you have that, you also need this”).
  • Groups 2–4 — business applications. Only apps with CIF = Yes are reviewed; “No” apps like DocuSign stay in the plan but are skipped.
  • The sequence starts with Full Responsibility of Providers (an assumption, no IRRP), then Break Glass (Azure, Zscaler), then Network — Break Glass must finish before Azure Network Cluster starts.
Last year’s lesson learned — act on this

Two groups was a mistake: too big, and the people at the end just waited. Feedback was that it was boring and people disengaged. The fix: four groups of 1.5 hours instead of two of three hours. Prisca had no recipe yet — she said explicitly to work the grouping out with Reto.

Prepare this before you lead it

You must explain Blocking / Semi-Blocking / Non-Blocking at the start so everyone understands the sequencing. The definitions sit in a collapsible section at the top of the FRCP sheet. Prisca still didn’t find them great — clarify anything unclear with Reto, and reuse her slides.

09 · Fact24 & Emergency Org

The lifeboat

Fact24 is what you use when you have nothing — no Confluence, no Jira, no SharePoint. It’s hosted externally on AWS, completely separate from Swiss Life systems, bought by the Holding group for crisis management. Login is single sign-on.

Fact24 & the ICT Emergency Organisation The lifeboat: what you use when Confluence, Jira and SharePoint are all gone. Why it exists Hosted externally on AWS, completely separate from Swiss Life systems. Bought by the Holding group for crisis management; ICT EO uses it too. If everything is down, this still works. Two modules SIM — the document store Admin → File Archive → ICT EO and SLAM ITSCM • Under ICT EO: templates for an emergency — checklists, agendas, journals • Under SLAM ITSCM: every IRRP + the Full Recovery Plan, as PDF • Plus a ZIP per app containing every linked document the owner would need offline ENS+ — the alerting world Alarms fire on three channels in parallel: SMS  +  email  +  phone call You choose how to dial in. Two alarm types are configured: Exercise Test Alarm — text says "Training" so nobody panics during a drill • The real one Who is in the ICT Emergency Organisation ICT EO Head 2 people: Head of Shared IT + Head of Security IT Owns the problem-solving Management Support Reto · Roland · Mark · YOU You manage the process, not the technical fix. Core Members Specialists who know their application. Only primaries; deputies step in if absent. Extended Pulled in when CISO, France, Germany or Luxembourg are affected. How an emergency call runs 1. Everyone joins the first call — that's the strategy. 2. The situation is explained to all. 3. The Head works through a checklist ("extended memory") so no point is forgotten, and decides who is needed. 4. Those not needed drop out and can be pulled back in. Things to know • Scope is IT situations only — building evacuation is a different department entirely. • Everyone with a role has triggered a test alarm once, so it never depends on one person. They rotate. • Prisca documented the whole setup on a Confluence page.
Diagram · Fact24’s two modules and the ICT Emergency Organisation
The rule to remember

Everything in Fact24 is uploaded manually. There is no synchronisation with Confluence or SharePoint. Every time an IRRP or the FRCP changes, the PDF has to be re-uploaded.

Your place in it

You are Management Support, alongside Reto, Roland and Mark — you manage the process, not the technical fix. The ICT EO Head (Head of Shared IT + Head of Security IT) owns the problem-solving and decides who is needed.

  • Everyone joins the first call — that’s the deliberate strategy. Then people drop out if they’re not needed, and can be pulled back in.
  • The Head works from a checklist — Prisca called it “extended memory” — so no point is forgotten.
  • Scope is IT situations only. Building evacuation and the general crisis staff are a different department entirely.
  • Alarms fire on three channels in parallel: SMS, email and phone call.
  • Everyone with a role has triggered a test alarm at least once, so it never depends on one person. They rotate.
  • Q3 exercise: Reto leads and is the Master. You’re Management Support with him and Roland.
Find this early

Prisca documented the entire Fact24 setup on a Confluence page — how she configured persons, groups, qualifications, alarms and templates, and how to trigger an exercise alarm. She noted she was “pretty much the only one who worked into this,” so locate that page in your first weeks.

10 · Your first 90 days

What to do, week by week

Everything below works backwards from 1 November. The full version lives in 06_My_Notes/First_90_Days_Playbook.md; the checklists in section 14 track it.

Weeks 1–2

Get switched on

Chase access (Jira/Xray, Confluence, SharePoint, the private EO channel, Fact24, ADONIS). Get swapped in for Prisca on the Two-Pager and on every Control Procedure. Meet Reto, Roland, Romano and Mark.

Weeks 3–4

Learn the live picture

Read the Epics Prisca annotated with status before leaving — that’s her written handover. Open the finished eFront report as your model. List the ~30 CIFs with owners and tiers. Read 3–4 IRRPs and the FRCP.

Early October

Prepare the cycle ⚠

Confirm the CIF list, create any missing IRRP, verify every Control Procedure points at the right two people, and book the December walkthrough dates now — people are booked out early.

Nov – Dec

Run it

Cross-check ~30 reviews, get approvals, upload each PDF to Fact24, then deliver the walkthrough in four groups and produce FRCP v3.0.

Ten things she’d tell you on day one

#Advice
1Get into the November cycle by early October. Everything else can wait; this can’t.
2You’re not the technical expert — and you’re not supposed to be.
3Never let new work land in November. November is for checking what exists.
4Persistence is the actual skill. Keep coming back; keep offering to take work off people.
5Write it down. Anything told to you verbally “floats somewhere in the air” — put it in the Epic.
6Use DORA as the lever. “We must, by law” unlocks time, money and cooperation.
7English by default. With English you’re always fine; with German you can lose someone.
8Everything in Fact24 is manual. Nothing syncs.
9Ask Romano for anything Jira or Control-Procedure related.
10Four small groups, not two big ones. People switch off in a three-hour walkthrough.
11 · Coverage & on-call

Who's reachable, and who gets called

Why "Best Effort" outside business hours is a deliberate, cost-based choice — and how you interface to the Emergency Organisation.

Coverage & the Emergency Organisation (the real picture) When people are reachable, who gets called, and why “Best Effort” outside hours is a deliberate choice. A day of coverage Trading hrs from ~07:00 Business / Support ~07:30 – 18:00 staffed, fast response Night & weekend “Best Effort” applies After ~22:00 little runs; an issue can often wait and be fixed before the next morning's market open. “Follow the sun” Branches sit in different time zones — Portugal, New York, Singapore, Hong Kong. During on-call a call can come from New York or Singapore. US trading hours especially matter for shared infrastructure. Who gets called (the interface to the Emergency Organisation) ITSCM / Test Manager holds the ITSCM topic; the interface Manager on Duty (MoD) ~4 people — coordinates the issue, escalates to the crisis team if needed Security: Support of the Week a named security contact each week contacts The on-call reality at Swiss Life AM • There is no formal paid 24/7 standby rota — it was set up once, then switched off for cost reasons (very few out-of-hours cases, and applications here are mostly not highly time-critical). • Instead, people are engaged if reachable; on weekends a willing person takes the time and responds. • At another firm she knew, new MoDs were hired who “had no idea how it works” — here, staff already know ~70–90%. How this connects to the Strategy This is exactly why the Strategy says RTO targets only hold during business hours — outside them, “Best Effort” applies. The Emergency Organisation is also tested in simulations, and the MoDs take part.
Diagram · Coverage hours, follow-the-sun, and the on-call reality
12 · Testing framework reference

Types → Methods → Modules → Packages

The formal structure from the AM 12.1.1 framework. You'll mostly work in theoretical (reviews / walkthroughs / questionnaires) and practical (single & multi tests) modes, with the big overall exercises once a year.

The Testing Framework — 4 Building Blocks They nest inside each other: a Type is done via a Method, applied to a Module, bundled into a Package. 1 · TYPE — the broad category Theoretical talk it through (paper) Practical do one piece for real Overall simulate the full disaster 2 · METHOD — the technique (6 of them) Review Walkthrough Single Test Multi Test Simulation Full Restart / Recovery Test 3 · MODULE — the "test object" (what you point a test at) IaaS Server Restart Database Recovery AD Restart DC Switch Over SaaS Restart …and more. Restart = bring a healthy system back · Recovery = restore lost/damaged data from backup. 4 · PACKAGE — a bundle of modules tested together Scheduled in the Test Activity Plans — short-term = 1 year, long-term = 3 years. Tests must vary the outage scenarios, especially S2 & S3 (data loss).
Diagram · The four building blocks of the testing framework
The DORA questionnaire (a real example)

For DORA, a set of ~25 services got a theoretical interview questionnaire that owners answered; responses were checked and completed by the 10 Jan 2025 deadline. It's all captured under one Epic — a good template for the theoretical type.

13 · Session notes

What each handover covered

Reconstructed from the recordings (rough Swiss-German transcription — treat confirm items as to-verify). Expand each.

  • Your role is an enabler, not a doer — six working principles (see section 01).
  • How a real test gets built, and the shortcut of documenting real events as tests (section 03).
  • The on-call reality: no paid 24/7 standby; "Best Effort" outside hours; follow-the-sun (section 07).
  • Tooling is Jira/Xray + Word reports; audits (DORA) need a complete, gap-free overview.
  • Logistics: Wednesday is the on-site day; Reto will introduce you.
  • She spoke High German so you could focus on the substance, not the dialect.
  • Toured Confluence (Operational Resilience page, Strategy PDF, framework, ADONIS processes, Emergency Org) — section 05.
  • Toured the Test Activity Plan (tiers, colours, countries, cadence, placeholders) — section 06.
  • Drilled into a live Epic → Story → Test Plan → Test Cases → Report, with the Bloomberg and data-centre examples — section 04.
  • Explained defects, sign-off, the dashboard, and yearly archiving in Jira.
  • Your scope is 50% ITSCM testing + 50% supporting Roland; the future SimCorp Dimension (SCD) SaaS project will reshape things.
  • Key dates: an event on 8 Aug; SCD release ~April; next session on the 22nd.
  • Toured the ITSCM Teams channel — mainly a filing place, plus a private EO channel only EO members can see. The Two-Pager is now verified and laminated copies go out physically.
  • Walked the Recovery Documentation: the IRRPs and the Full Recovery Plan, plus archive folders for decommissioned and no-longer-CIF apps.
  • Explained the annual IRRP review driven by Control Procedures in Jira — 1 Nov to 30 Nov, then the walkthrough in week 1 of December (sections 07–08).
  • Showed Fact24: the SIM document store and the ENS+ alerting world, and the whole ICT Emergency Organisation structure (section 09).
  • Confirmed all EO communication is in English; the company language is English.
  • Background: the SimCorp release is running suboptimally (first time on the newest release, needed for the on-prem → SaaS move), targeting 8 August go-live. Roland finishes it after she leaves; next release maybe May/June 2027.
  • She is documenting the status of every running test inside the Epics before she goes, so it’s accessible to Reto too. The eFront test and report are complete.
  • She’ll ask Reto whether Mark Zwigart (In and Out), the external who ran this before her, can support your first cycle.

Got the next recording? Add it here — send it over and this section (and the checklists) will be extended.

14 · Checklists

Track your readiness

Tick items off as you complete them — progress saves in this browser and feeds the sidebar meter. HOT marks the highest-priority items.

First weeks — access & setup

0%

Create one test, end-to-end

0%

The Nov–Dec cycle (your biggest deadline)

0%

Ask in session 3 (the 22nd)

0%

30-day readiness (after she leaves)

0%
15 · Glossary, people & dates

Quick reference

Terms you'll hear daily

TermMeaning
Epic / Story / Test Plan / Test Case / ReportThe five building blocks of every test in Jira/Xray (section 04).
XrayThe Jira plug-in for test management (Test Plans, Test Cases, Execute).
Defect / BugA logged failure, copied from a test case, with a criticality.
Start vs. RecoverStart = bring a healthy system back (S1). Recover = restore from backup (S2/S3).
Full Recovery (FRCP)The yearly rebuild-everything-in-order exercise.
UATTest/acceptance environment — where you run tests without touching production.
ADONISSwiss Life's process-modelling tool (roles, responsibilities, the Emergency Org process).
Tier 1 / 2 / 3Application criticality bands that drive test frequency.
KVGA separate German Swiss Life company (a country/entity column in the plan).
MoDManager on Duty — ~4 people you contact for out-of-hours issues.
RTO / RPOMax downtime / max data loss (section 02).
S1 / S2 / S3No data loss / data loss / ransomware.
O1 / O2 / O3Geo-redundant sync / geo-redundant backup / immutable backup.
DORA / FINMAEU regulation / Swiss regulator — both require tested, provable recovery.
CIFCritical Important Function — an app critical enough to need an IRRP and an annual review (~30 of them).
Control ProcedureThe Jira mechanism that fires the annual IRRP review on 1 November and records it as audit evidence.
ICSThe internal control framework. One ICS control covers the IRP + FRCP; all Control Procedures hang off it.
Fact24External AWS-hosted platform: document store (SIM) + alerting (ENS+) for when everything else is down.
Two-PagerThe laminated contact/role card for the ICT Emergency Organisation. You’re on it as Management Support.
Break GlassEmergency-access accounts (Azure, Zscaler) — the true first step of the Full Recovery sequence.
Blocking / Semi- / Non-BlockingWhether a recovery step must finish before the next can start. You must be able to explain these.
Recovery ClassThe ordering bands in the FRCP, starting at Class 10 (IT basis infrastructure).
x.9 / x.0IRRP versions: x.9 = reviewed by the owner, awaiting approval; x.0 = approved and final.
InventxThe provider for the St. Gallen / Chur data-centre tests.
SimCorp Dimension (SCD)Major SaaS project reshaping future testing; Roland leads its test management.

People & culture

Who's who confirm names

Your circle

Prisca — your predecessor. Reto — ITSCM Manager: approves IRRPs, owns strategy, leads the Q3 exercise. Romano (Atlassian team) — built the Control Procedures; ask him for anything Jira. Mark Zwigart (In and Out) — the external who may support your first cycle. Reto — your predecessor. Reto — ITSCM Manager: approves IRRPs, owns strategy, leads the Q3 exercise. Roland — you support his role 50%; he finishes the SimCorp release. Romano (Atlassian team) — built the Control Procedures; ask him for anything Jira. Mark — fellow Management Support in the EO. Mark Zwigart (In and Out) — the external who ran this before Prisca; may support your first cycle.

Roland — you support his role 50%; he keeps the journal in exercises, knows the process. App owners / technical contacts — the specialists you enable.

Ways of working

Mostly remote

Wednesday is the on-site day (Reto & Roland usually in). ~95% of meetings are Teams chat/calls; even a Security CAB you might lead is remote. Go say hi in person when you can — relationships build over time (and coffee).

Dates to hold

WhenWhat
1 SepOfficial handover info starts arriving on your Swiss Life email.
The 22ndNext handover session (the "Full" walkthrough + Emergency Org with Reto). confirm
8 AugAn event she mentioned; after it, fewer sessions this year. confirm
Q3First exercise you'll take part in.
Early OctPrepare the annual cycle — scoping, Control Procedures, book walkthrough dates. Non-negotiable.
1–30 NovIRRP review window — ~30 Control Procedures, four weeks.
Week 1 DecFull Recovery Plan walkthrough — four groups of ~1.5h.
Mid-Jan – mid-FebContact all owners to lock the next year’s test plan.
~AprilSimCorp Dimension (SCD) SaaS release. confirm
16 · Study plan

40 days to get ready

One focused session a day, about 30 minutes, from 23 July to 31 August. Import the calendar file for a daily reminder at 19:00, and tick each day off here as you go.

The principle

Every session has one goal, and the goal is always “be able to explain this without looking” — not “have read it”. Reading feels productive; explaining is what sticks. If you can’t say it out loud at the end of a session, repeat that session rather than moving on.

Sundays are deliberately lighter — review and consolidation, no new material.

Import it

Mac

Double-click ITSCM_Study_Plan.ics. Add it to a new calendar called ITSCM so you can hide or delete the whole thing later in one go.

Import it

iPhone / Google

iPhone: email the file to yourself and tap it. Google Calendar: Settings → Import & export → Import.

Adjust it

Different time?

Events are set to 19:00 with a 10-minute heads-up. Select them all in your calendar and drag, or ask for a regenerated file at another hour.

If you miss days

Don’t binge to catch up. Continue from where you are — week 6 is consolidation and can be shortened without losing anything important.

Week 1 — Foundations

0%

Week 2 — The source documents

0%

Week 3 — The role & the workflow

0%

Week 4 — The annual cycle

0%

Week 5 — Emergency & governance

0%

Week 6 — Consolidation

0%