Walltime exceeded time out in atmos_main

Hello,

I am running two suites (u-eb430 and u-eb461), both copies of u-dr799 (free-running AMIP configuration). They both keep failing with walltime exceeded error. The first month runs in just over an hour, the second month takes over 2 hours (so I increased wallclock time to 3 hours), and now the third month is exceeding 3 hours and failing.

I am not sure why each month is taking longer to run and how to fix this.

Thank you!

Hi Isangha,

It would be useful to check where the extra time is being spent before increasing the walltime further.

Could you please provide:

-The atmos_main timing information from the first successful month and the failed cycle so we know which is taking longer and error logs

-The suite cycling configuration (cycle length and restart/dump frequency).

-Any changes made compared with the original u-dr799 suite (especially diagnostics/output settings).

Best,

Juan

Hi Juan,

Thank you for your response.

The atmos_main timing info from the job.out file for the successful month (Jan) :

============================= PBS epilogue =============================

End of Job Report
Run at 2026-08-10 10:25:38 for job 9857007.ehz200
Submitted: 2026-08-10 09:24:07
Queued: 2026-08-10 09:24:07
Started: 2026-08-10 09:24:11
Completed: 2026-08-10 10:25:37
Processed: 2026-08-10 10:25:38
Queued Time: 0:00:04
Elapsed Time: 1:01:26 (3686 seconds, 34% of total)
Walltime Limit: 3:00:00
Job Name: atmos_main.19790101T0000Z.u-eb430-run1
Job Queue: collab
Owner: isabelle.sangha.ext
Project: other
Funding: unknown
Output: ln12:/lustre/ehz2col/collaboration/home/users/isabelle.sangha.ext/cylc-run/u-eb430/run1/log/job/19790101T0000Z/atmos_main/01/job.out
Error: ln12:/lustre/ehz2col/collaboration/home/users/isabelle.sangha.ext/cylc-run/u-eb430/run1/log/job/19790101T0000Z/atmos_main/01/job.err
Job Directory: /home/users/isabelle.sangha.ext/pbs.9857007.ehz200.x8z
Total Nodes: 3 (coretype: milan)
Total Tasks: 768
Memory Used: 67.6GB of 711.0GB (10% of total)
Total CPU Time: 470.4ks
Primary Node: nide1644
Run Version: 1
Exit Status: 0 (Job execution was successful)

and successful month (Feb) :

============================= PBS epilogue =============================

End of Job Report
Run at 2026-08-10 13:09:32 for job 9873697.ehz200
Submitted: 2026-08-10 10:25:41
Queued: 2026-08-10 10:25:41
Started: 2026-08-10 11:06:04
Completed: 2026-08-10 13:09:31
Processed: 2026-08-10 13:09:32
Queued Time: 0:40:23
Elapsed Time: 2:03:27 (7407 seconds, 69% of total)
Walltime Limit: 3:00:00
Job Name: atmos_main.19790201T0000Z.u-eb430-run1
Job Queue: collab
Owner: isabelle.sangha.ext
Project: other
Funding: unknown
Output: ln10:/lustre/ehz2col/collaboration/home/users/isabelle.sangha.ext/cylc-run/u-eb430/run1/log/job/19790201T0000Z/atmos_main/01/job.out
Error: ln10:/lustre/ehz2col/collaboration/home/users/isabelle.sangha.ext/cylc-run/u-eb430/run1/log/job/19790201T0000Z/atmos_main/01/job.err
Job Directory: /home/users/isabelle.sangha.ext/pbs.9873697.ehz200.x8z
Total Nodes: 3 (coretype: milan)
Total Tasks: 768
Memory Used: 74.1GB of 711.0GB (10% of total)
Total CPU Time: 946.5ks
Primary Node: nide1652
Run Version: 1
Exit Status: 0 (Job execution was successful)

and failed month (mar) :

Atm_Step: Timestep 6357 Model time: 1979-03-29 07:00:00
============================= PBS epilogue =============================

End of Job Report
Run at 2026-08-10 16:19:38 for job 9923178.ehz200
Submitted: 2026-08-10 13:09:35
Queued: 2026-08-10 13:09:35
Started: 2026-08-10 13:18:27
Completed: 2026-08-10 16:19:37
Processed: 2026-08-10 16:19:38
Queued Time: 0:08:52
Elapsed Time: 3:01:10 (10870 seconds, 101% of total)
Walltime Limit: 3:00:00
Job Name: atmos_main.19790301T0000Z.u-eb430-run1
Job Queue: collab
Owner: isabelle.sangha.ext
Project: other
Funding: unknown
Output: ln12:/lustre/ehz2col/collaboration/home/users/isabelle.sangha.ext/cylc-run/u-eb430/run1/log/job/19790301T0000Z/atmos_main/01/job.out
Error: ln12:/lustre/ehz2col/collaboration/home/users/isabelle.sangha.ext/cylc-run/u-eb430/run1/log/job/19790301T0000Z/atmos_main/01/job.err
Job Directory: /home/users/isabelle.sangha.ext/pbs.9923178.ehz200.x8z
Total Nodes: 3 (coretype: milan)
Total Tasks: 768
Memory Used: 73.8GB of 711.0GB (10% of total)
Total CPU Time: 1.4Ms
Primary Node: nide1414
Run Version: 1
Exit Status: -29 (Job exec failed due to exceeding walltime)

The suite cycling configuration:

cycling frequency : P1M

arch_dump_freq : yearly

Changes made with original u-dr799 (Diff [363158:363356] for e/b/4/3/0 – roses-u):

  • added UM and UKCA branches (both have been run successfully in AMIP nudged suite)
  • added some daily and monthly STASH output (less additional STASH than what was included in my successful AMIP nudged suite with the same branches)

Thank you for your help. Please let me know if I can provide any more information to help get this suites running.

Best,
Isabelle

Hi Isabelle,

Thanks for providing the additional information.

This suggests that the extra runtime is occurring during the model execution rather than only during queueing or output processing.

Can you please share the UM timing information from atmos_main for one of the faster cycles and the slower cycle (March)? The timing summary in the pe_output file would be useful.

This should help identify whether the additional time is coming from a particular component like UKCA, physics, dynamics or STASH/output.

Also, could you confirm whether the UKCA configuration is exactly the same as in your successful one or whether there are any differences?

Thanks,
Juan

Hi Juan,

After digging in the job.out file it appears that each run is starting from Jan 1 rather than the first of the month. For example, cylc-run/u-eb430/run1/log/job/19790301T0000Z/atmos_main/01/job.out has timesteps:
Atm_Step: Timestep 2 Model time: 1979-01-01 00:40:00
Atm_Step: Timestep 3 Model time: 1979-01-01 01:00:00
Atm_Step: Timestep 4 Model time: 1979-01-01 01:20:00
Atm_Step: Timestep 5 Model time: 1979-01-01 01:40:00
Atm_Step: Timestep 6 Model time: 1979-01-01 02:00:00
Atm_Step: Timestep 7 Model time: 1979-01-01 02:20:00
Atm_Step: Timestep 8 Model time: 1979-01-01 02:40:00
Atm_Step: Timestep 9 Model time: 1979-01-01 03:00:00
Atm_Step: Timestep 10 Model time: 1979-01-01 03:20:00
Atm_Step: Timestep 11 Model time: 1979-01-01 03:40:00

I had changed the rose suite setting for cycling frequency from P3M (u-dr799) to P1M (u-eb430), but does this mean that it may still be doing something on a 3 month frequency so it is restarting from 3 months earlier?

Thanks,
Isabelle

Hi Isabelle,

Changing the Cylc cycle frequency from P3M to P1M changes the scheduling frequency but it may not automatically change the UM run length or restart selection.

Can you please check the restart/run-length settings in u-eb430 compared with u-dr799?

what restart dump is selected at the start of the 19790301T0000Z cycle?

Thanks,
Juan

FYI there appears to be some kind of outage to Monsoon access right now. I’ve sent a Teams message to the maintainers.

Cheers

Jonny

The dump frequency was 90 days, which I have just changed to 30 days to match the EXPT_RESUB=‘P1M’ and am rerunning.

The fun-length was changed from 3 months in u-dr799 to 39 years in u-eb430.

I am not sure how to see what restart dump is selected at the start of the 19790301T0000Z cycle, which file would this be in?

Thanks,

Isabelle

Hi Isabelle,

Thanks for checking this.

Let see if the dump frequency to 30 days to match the monthly resubmission should resolve this.

To check which restart dump is being used, can you check the atmos_main job.out for the March cycle?

path_to/19790301T0000Z/atmos_main/01/job.out

Search for dump, restart, or start messages. Let me know how it goes.

Best regards,

Juan

Hi Juan,

Changing the dump frequency to 30 days did resolve the issue. Thank you for all your help with this!

Hi Isabelle,

Good to know that it is now working!

Best regards,
Juan