Perturbed start dump not changing run outputs

Hello,

I’m a bit stumped on this and am hoping someone can find something I’ve overlooked! I am trying to set up a 3 member ensemble of AMIP historical free-running simulations from 1979 - 2014 (copies of u-dr799). My three members are run in suites :

  • u-eb430 (using the default start dump)
  • u-eb974 (using a perturbed start dump /projects/ukca-cam/isabelle.sangha.ext/start_dumps/by791a.da19790101_00_perturb1)
  • u-eb982 (using a perturbed start dump /projects/ukca-cam/isabelle.sangha.ext/start_dumps/by791a.da19790101_00_perturb2)

The perturbed start dumps are made using perturb_theta and seem to be different (see mule-cumf --summary output below) :
%%%%%%%%%%%%%%%%%%%%%%%%%%%%%

  • CUMF-II Comparison Report *
    %%%%%%%%%%%%%%%%%%%%%%%%%%%%%

File 1: /projects/ukca-cam/isabelle.sangha.ext/start_dumps/by791a.da19790101_00_perturb1
File 2: /projects/ukca-cam/isabelle.sangha.ext/start_dumps/by791a.da19790101_00_perturb2
Files DO NOT compare

  • 0 differences in fixed_length_header (with 7 ignored indices)
  • 86 field differences, of which 86 are in data

Compared 55379/55379 fields, with 55293 matches

Maximum RMS diff as % of data in file 1: 1.5242971721894078e-14 (field 4151)
Maximum RMS diff as % of data in file 2: 1.5242971721894078e-14 (field 4151)

The same is true for the astart files:

%%%%%%%%%%%%%%%%%%%%%%%%%%%%%
* CUMF-II Comparison Report *
%%%%%%%%%%%%%%%%%%%%%%%%%%%%%

File 1: cylc-run/u-eb974/run1/share/data/eb974a.astart
File 2: cylc-run/u-eb982/run1/share/data/eb982a.astart
Files DO NOT compare
* 0 differences in fixed_length_header (with 7 ignored indices)
* 86 field differences, of which 86 are in data

Compared 17844/17844 fields, with 17758 matches

Maximum RMS diff as % of data in file 1: 1.5242971721894078e-14 (field 4065)
Maximum RMS diff as % of data in file 2: 1.5242971721894078e-14 (field 4065

But when I compare dump files from a few years after the start of the simulations there don’t seem to be many (if any) differences. (10/11 of the fields that have differences are NaN) :

%%%%%%%%%%%%%%%%%%%%%%%%%%%%%

  • CUMF-II Comparison Report *
    %%%%%%%%%%%%%%%%%%%%%%%%%%%%%

File 1: check_ens_diff/eb974a.da19850101_00
File 2: check_ens_diff/eb982a.da19850101_00
Files DO NOT compare

  • 0 differences in fixed_length_header (with 7 ignored indices)
  • 11 field differences, of which 11 are in data

Compared 45105/45105 fields, with 45094 matches

Maximum RMS diff as % of data in file 1: 222.93860914299302 (field 43964)
Maximum RMS diff as % of data in file 2: 111.39620408658773 (field 43964)

Thanks for any help with figuring out why the different start dumps aren’t producing many differences in the model fields.

Hi Isabelle.

I’m trying to compare some other files from these runs just to check the evolution of some specific variables but it looks like there is some postprocessing going on which is moving or deleting files (e.g. *a.pa*)…

$ ls u-eb974/run1/share/data/History_Data/*pa*
u-eb974/run1/share/data/History_Data/eb974a.pa1988aug 
u-eb974/run1/share/data/History_Data/eb974a.pa1988sep

$ ls u-eb982/run1/share/data/History_Data/*pa*
u-eb982/run1/share/data/History_Data/eb982a.pa1988jul  
u-eb982/run1/share/data/History_Data/eb982a.pa1988jun

Do you know where these files are being moved to?

Cheers

Jonny

Hi Johnny,

I don’t know why those files would be moved. Is there a rose setting that moves/deletes files from History_Data?

Thanks,

Izzy

No problem. I’ll do a short test run now to see if I can reproduce what you’re seeing. Any differences from a perturbed restart file will show up very quickly anyway. :smiley:

Cheers

Hi Isabelle.

Could you please let me know how you are accessing mule on Monsoon3? It’s a requirement of perturb_theta.py. Are you loading a conda environment or similar?

Thanks

Jonny

Hi Jonny,

I do:

module load scitools

export PYTHONPATH=/projects/metoff/umdir/mule/mule-2025.10.1/python3.12.12_numpy-2.3.5/no-openmp/lib

module load um_tools

Hi Isabelle.

I’ve done a 1 month run of each of u-eb430, u-eb974, and u-eb982 I copied these exactly from your ~isabelle.sangha.ext/roses/u-.... directories.

I do actually see marked changes in the results of the three runs and indeed u-eb982 seems to be heading towards a crash with some unphysical values of, e.g., surface temperature…

I’ve also noticed that you have not used ‘peturb1’ for u-eb974. This is a screenshot showing the difference ainitial values used in the three runs…

I suggest you try running some fresh 1 month jobs to confirm that you get the same results as me, i.e., that the restarts don’t compare.

Incidentally, using the same method as you mentioned (here)[Perturbed start dump not changing run outputs - #6 by isangha], I get the following errors when I run perturb_theta.py

$ python perturb_theta.py --output ./by791a.da19790101_00_perturb1-jonny /projects/ukesm/harley.kelly.mon/restarts/u-by791/by791a.da19790101_00

Perturbing Field - Sec 0, Item 388: ThetaVD After Timestep

[snip]

[ERROR] Mule failed to validate fields writing to file:
./by791a.da19790101_00_perturb1-jonny
Traceback (most recent call last):
  File "/lustre/ehz2col/collaboration/home/users/jonny.williams.ext/perturb_theta/perturb_theta.py", line 166, in main
    dfile.to_file(dump_out)
  File "/projects/metoff/umdir/mule/mule-2025.10.1/python-3.12.12_numpy-2.3.5/openmp/lib/mule/__init__.py", line 1406, in to_file
    self.validate(filename=output_file_or_path)
  File "/projects/metoff/umdir/mule/mule-2025.10.1/python-3.12.12_numpy-2.3.5/openmp/lib/mule/validators.py", line 200, in validate_umf
    raise ValidateError(filename, "\n".join(validation_errors))
mule.validators.ValidateError: Failed to validate
File: ./by791a.da19790101_00_perturb1-jonny
Field validation failures:
  Fields (727,728,729,730,773, ... 11 total fields)
Field grid longitudes inconsistent
  File grid : 0.0 to 358.125, spacing 1.875
  Field grid: 0.5 to 359.5, spacing 1.0
  Extents should be within 1 field grid-spacing

During handling of the above exception, another exception occurred:

Traceback (most recent call last):
  File "/lustre/ehz2col/collaboration/home/users/jonny.williams.ext/perturb_theta/perturb_theta.py", line 182, in <module>
    main()
  File "/lustre/ehz2col/collaboration/home/users/jonny.williams.ext/perturb_theta/perturb_theta.py", line 178, in main
    raise mule.validators.ValidateError(dump_in, '')
mule.validators.ValidateError: Failed to validate
File: /projects/ukesm/harley.kelly.mon/restarts/u-by791/by791a.da19790101_00

login11|Fri Sep 04|14:22:20 UTC|perturb_theta$

Cheers

Jonny

Hi again

So for some reason perturb_theta.py works on ARCHER2 but not Monsoon3 for me. I’m doing a test now (on Monsoon3) using restarts generated from ARCHER2.

I’ll let you know on Monday how it’s gotten on.

Cheers

Jonny

Hi Jonny,

Thanks for looking into this. Are these runs done on Archer2 or Monsoon3? Mohit advised me to add this line :

mule.stashmaster.STASHMASTER_PATH_PATTERN = ‘/projects/metoff/umdir/vn13.9/ctldata/STASHmaster/STASHmaster_A’

at the start of the perturb_theta.py to get it to run on Monsoon3.

The different ainitial (eb430a.da19800101_00_1979) is from today because I am now trying to perturb the runs by using dump file from a year after the start, rather than perturb_theta.

It’s weird the large differences u-eb982 because my initial run for u-eb982 using the perturb2 start dump ran completely fine (had made it almost 10 years when I stopped it), but all the fields were the same as u-eb430.

Hi there.

Are these runs done on Archer2 or Monsoon3?

I’m doing the runs on Monsoon3 (under ~jonny.williams.ext). I only used ARCHER2 to create some perturbed restarts.

Mohit advised me to add this line :

mule.stashmaster.STASHMASTER_PATH_PATTERN = ‘/projects/metoff/umdir/vn13.9/ctldata/STASHmaster/STASHmaster_A’

Thank you, that’s worked for me.

The different ainitial (eb430a.da19800101_00_1979) is from today because I am now trying to perturb the runs by using dump file from a year after the start, rather than perturb_theta.

I think this is a good idea; certainly worth a try.

I’m trying a version with the reconfiguration turned off to see if this makes any difference. The results I get when running with my newly perturbed restart files do differ according to mule-cumf but they are extremely small (in agreement with your previous message). According to the ‘docstring’ in perturb_theta.py, the script is designed to fix ‘grid point storms’ which occur mid-run, not at the start. So it’s possible that the reconfiguration is somehow ironing out the differences intruduced by perturb_theta.py.

Another way of doing a qualitatively similar thing to perturbing theta (ostensibly to fix those ‘grid point storms’ I mentioned) is to change the number of calls to convection for 1 month to tweak the integration of the model, e.g. see here.

Yet another possibility is to just use some start dumps other years in the same workflow as the standard one from u-by791. These are available on MASS…

$ moo ls moose:/crum/u-by791/ada.file|grep -C 1 1979
moose:/crum/u-by791/ada.file/by791a.da19780101_00
moose:/crum/u-by791/ada.file/by791a.da19790101_00
moose:/crum/u-by791/ada.file/by791a.da19800101_00

To use these you’d need to overwrite the start year in the dump but that’s straightforward using the rose GUI.

All the best.

Jonny

Hi @isangha.

Couple of quick updates.

  1. Turning off the reconfiguration does not work. It fails with a ‘wrong number of fields found’ error which isn’t entirely surprising.

  2. Changing the number of calls to convection for the 1st month works for me and gives differing results which are easily visible by eye (2 calls to convection on left; 3 on right) for the 1st of February 1979 restart. :partying_face:

If you want to go down this route then this is the recipe for doing do.

  1. Edit app/um/rose-app.conf in ~/roses/[workflow ID] so that n_conv_calls=3 (the default is 2).
  2. Launch the workflow
  3. Hold the 2nd month so that it doesn’t run automatically when the 1st month has completed. You can do this by right-clocking on the 2nd month’s atmos_main task in the Cylc web UI.
  4. Change n_conv_calls bask to 2 in ~/roses/[workflow ID].
  5. Run cylc vr [workflow ID]/[run ID] (e.g. cylc vr u-ab123/run1).
  6. Release the previously held task, again by right-clicking.

This is admittedly a little bit fiddly but it’s a useful way of manipulating one’s workflows at runtime if you don’t want to use different restarts from MASS for example (see above).

Cheers.

Jonny

Hi Jonny,

Thank you for all your help with this! The runs I have set off using a dump file from one year into the simulation and then changing the year and using this as the start dump for the next ensemble member seems to be running, and a check of monthly output from a year it looks to be different (based on mule-cumf).

If this ends up failing, I will use the method you’ve outlined and change the number of convection calls!