# Walltime exceeded time out in atmos\_main

**URL:** <https://cms-helpdesk.ncas.ac.uk/t/walltime-exceeded-time-out-in-atmos-main/2144>\
**Category:** Unified Model\
**Created:** [11 August 2026 08:04 UTC](https://cms-helpdesk.ncas.ac.uk/t/walltime-exceeded-time-out-in-atmos-main/2144 "2026-08-11T08:04:49Z")\
**Posts on this page:** 12\
**Page:** 1

<div class="post-metadata">

**Author:** ![isangha](https://avatars.discourse-cdn.com/v4/letter/i/82dd89/32.png) [@isangha](https://cms-helpdesk.ncas.ac.uk/u/isangha)\
**Post date:** [11 August 2026 08:04 UTC](https://cms-helpdesk.ncas.ac.uk/t/walltime-exceeded-time-out-in-atmos-main/2144/1 "2026-08-11T08:04:49Z")

</div>

Hello,

I am running two suites (u-eb430 and u-eb461), both copies of u-dr799 (free-running AMIP configuration). They both keep failing with walltime exceeded error. The first month runs in just over an hour, the second month takes over 2 hours (so I increased wallclock time to 3 hours), and now the third month is exceeding 3 hours and failing.

I am not sure why each month is taking longer to run and how to fix this.

Thank you!

---

<div class="post-metadata">

**Author:** ![Juan\_Bilbao](https://dub1.discourse-cdn.com/flex013/user_avatar/cms-helpdesk.ncas.ac.uk/juan_bilbao/32/458_2.png) [@Juan\_Bilbao](https://cms-helpdesk.ncas.ac.uk/u/Juan_Bilbao)\
**Post date:** [11 August 2026 08:54 UTC](https://cms-helpdesk.ncas.ac.uk/t/walltime-exceeded-time-out-in-atmos-main/2144/2 "2026-08-11T08:54:24Z")

</div>

Hi Isangha,

It would be useful to check where the extra time is being spent before increasing the walltime further.

Could you please provide:

-The `atmos_main` timing information from the first successful month and the failed cycle so we know which is taking longer and error logs

-The suite cycling configuration (cycle length and restart/dump frequency).

-Any changes made compared with the original `u-dr799` suite (especially diagnostics/output settings).

Best,

Juan

---

<div class="post-metadata">

**Author:** ![isangha](https://avatars.discourse-cdn.com/v4/letter/i/82dd89/32.png) [@isangha](https://cms-helpdesk.ncas.ac.uk/u/isangha)\
**Post date:** [11 August 2026 09:05 UTC](https://cms-helpdesk.ncas.ac.uk/t/walltime-exceeded-time-out-in-atmos-main/2144/3 "2026-08-11T09:05:47Z")

</div>

Hi Juan,

Thank you for your response.

The atmos\_main timing info from the job.out file for the successful month (Jan) :

============================= PBS epilogue =============================

End of Job Report  
Run at 2026-08-10 10:25:38 for job 9857007.ehz200  
Submitted: 2026-08-10 09:24:07  
Queued: 2026-08-10 09:24:07  
Started: 2026-08-10 09:24:11  
Completed: 2026-08-10 10:25:37  
Processed: 2026-08-10 10:25:38  
Queued Time: 0:00:04  
Elapsed Time: 1:01:26 (3686 seconds, 34% of total)  
Walltime Limit: 3:00:00  
Job Name: atmos\_main.19790101T0000Z.u-eb430-run1  
Job Queue: collab  
Owner: isabelle.sangha.ext  
Project: other  
Funding: unknown  
Output: ln12:/lustre/ehz2col/collaboration/home/users/isabelle.sangha.ext/cylc-run/u-eb430/run1/log/job/19790101T0000Z/atmos\_main/01/job.out  
Error: ln12:/lustre/ehz2col/collaboration/home/users/isabelle.sangha.ext/cylc-run/u-eb430/run1/log/job/19790101T0000Z/atmos\_main/01/job.err  
Job Directory: /home/users/isabelle.sangha.ext/pbs.9857007.ehz200.x8z  
Total Nodes: 3 (coretype: milan)  
Total Tasks: 768  
Memory Used: 67.6GB of 711.0GB (10% of total)  
Total CPU Time: 470.4ks  
Primary Node: nide1644  
Run Version: 1  
Exit Status: 0 (Job execution was successful)

and successful month (Feb) :

============================= PBS epilogue =============================

End of Job Report  
Run at 2026-08-10 13:09:32 for job 9873697.ehz200  
Submitted: 2026-08-10 10:25:41  
Queued: 2026-08-10 10:25:41  
Started: 2026-08-10 11:06:04  
Completed: 2026-08-10 13:09:31  
Processed: 2026-08-10 13:09:32  
Queued Time: 0:40:23  
Elapsed Time: 2:03:27 (7407 seconds, 69% of total)  
Walltime Limit: 3:00:00  
Job Name: atmos\_main.19790201T0000Z.u-eb430-run1  
Job Queue: collab  
Owner: isabelle.sangha.ext  
Project: other  
Funding: unknown  
Output: ln10:/lustre/ehz2col/collaboration/home/users/isabelle.sangha.ext/cylc-run/u-eb430/run1/log/job/19790201T0000Z/atmos\_main/01/job.out  
Error: ln10:/lustre/ehz2col/collaboration/home/users/isabelle.sangha.ext/cylc-run/u-eb430/run1/log/job/19790201T0000Z/atmos\_main/01/job.err  
Job Directory: /home/users/isabelle.sangha.ext/pbs.9873697.ehz200.x8z  
Total Nodes: 3 (coretype: milan)  
Total Tasks: 768  
Memory Used: 74.1GB of 711.0GB (10% of total)  
Total CPU Time: 946.5ks  
Primary Node: nide1652  
Run Version: 1  
Exit Status: 0 (Job execution was successful)

and failed month (mar) :

Atm\_Step: Timestep 6357 Model time: 1979-03-29 07:00:00  
============================= PBS epilogue =============================

End of Job Report  
Run at 2026-08-10 16:19:38 for job 9923178.ehz200  
Submitted: 2026-08-10 13:09:35  
Queued: 2026-08-10 13:09:35  
Started: 2026-08-10 13:18:27  
Completed: 2026-08-10 16:19:37  
Processed: 2026-08-10 16:19:38  
Queued Time: 0:08:52  
Elapsed Time: 3:01:10 (10870 seconds, 101% of total)  
Walltime Limit: 3:00:00  
Job Name: atmos\_main.19790301T0000Z.u-eb430-run1  
Job Queue: collab  
Owner: isabelle.sangha.ext  
Project: other  
Funding: unknown  
Output: ln12:/lustre/ehz2col/collaboration/home/users/isabelle.sangha.ext/cylc-run/u-eb430/run1/log/job/19790301T0000Z/atmos\_main/01/job.out  
Error: ln12:/lustre/ehz2col/collaboration/home/users/isabelle.sangha.ext/cylc-run/u-eb430/run1/log/job/19790301T0000Z/atmos\_main/01/job.err  
Job Directory: /home/users/isabelle.sangha.ext/pbs.9923178.ehz200.x8z  
Total Nodes: 3 (coretype: milan)  
Total Tasks: 768  
Memory Used: 73.8GB of 711.0GB (10% of total)  
Total CPU Time: 1.4Ms  
Primary Node: nide1414  
Run Version: 1  
Exit Status: -29 (Job exec failed due to exceeding walltime)

The suite cycling configuration:

cycling frequency : P1M

arch\_dump\_freq : yearly

Changes made with original u-dr799 ([Diff [363158:363356] for e/b/4/3/0 – roses-u](https://code.metoffice.gov.uk/trac/roses-u/changeset?reponame=&new=363356%40e%2Fb%2F4%2F3%2F0&old=363158%40e%2Fb%2F4%2F3%2F0)):

- added UM and UKCA branches (both have been run successfully in AMIP nudged suite)
- added some daily and monthly STASH output (less additional STASH than what was included in my successful AMIP nudged suite with the same branches)

Thank you for your help. Please let me know if I can provide any more information to help get this suites running.

Best,  
Isabelle

---

<div class="post-metadata">

**Author:** ![Juan\_Bilbao](https://dub1.discourse-cdn.com/flex013/user_avatar/cms-helpdesk.ncas.ac.uk/juan_bilbao/32/458_2.png) [@Juan\_Bilbao](https://cms-helpdesk.ncas.ac.uk/u/Juan_Bilbao)\
**Post date:** [11 August 2026 10:05 UTC](https://cms-helpdesk.ncas.ac.uk/t/walltime-exceeded-time-out-in-atmos-main/2144/4 "2026-08-11T10:05:12Z")

</div>

Hi Isabelle,

Thanks for providing the additional information.

This suggests that the extra runtime is occurring during the model execution rather than only during queueing or output processing.

Can you please share the UM timing information from `atmos_main` for one of the faster cycles and the slower cycle (March)? The timing summary in the `pe_output` file would be useful.

This should help identify whether the additional time is coming from a particular component like UKCA, physics, dynamics or STASH/output.

Also, could you confirm whether the UKCA configuration is exactly the same as in your successful one or whether there are any differences?

Thanks,  
Juan

---

<div class="post-metadata">

**Author:** ![isangha](https://avatars.discourse-cdn.com/v4/letter/i/82dd89/32.png) [@isangha](https://cms-helpdesk.ncas.ac.uk/u/isangha)\
**Post date:** [11 August 2026 10:08 UTC](https://cms-helpdesk.ncas.ac.uk/t/walltime-exceeded-time-out-in-atmos-main/2144/5 "2026-08-11T10:08:14Z")

</div>

Hi Juan,

After digging in the job.out file it appears that each run is starting from Jan 1 rather than the first of the month. For example, cylc-run/u-eb430/run1/log/job/19790301T0000Z/atmos\_main/01/job.out has timesteps:  
Atm\_Step: Timestep 2 Model time: 1979-01-01 00:40:00  
Atm\_Step: Timestep 3 Model time: 1979-01-01 01:00:00  
Atm\_Step: Timestep 4 Model time: 1979-01-01 01:20:00  
Atm\_Step: Timestep 5 Model time: 1979-01-01 01:40:00  
Atm\_Step: Timestep 6 Model time: 1979-01-01 02:00:00  
Atm\_Step: Timestep 7 Model time: 1979-01-01 02:20:00  
Atm\_Step: Timestep 8 Model time: 1979-01-01 02:40:00  
Atm\_Step: Timestep 9 Model time: 1979-01-01 03:00:00  
Atm\_Step: Timestep 10 Model time: 1979-01-01 03:20:00  
Atm\_Step: Timestep 11 Model time: 1979-01-01 03:40:00

I had changed the rose suite setting for cycling frequency from P3M (u-dr799) to P1M (u-eb430), but does this mean that it may still be doing something on a 3 month frequency so it is restarting from 3 months earlier?

Thanks,  
Isabelle

---

<div class="post-metadata">

**Author:** ![Juan\_Bilbao](https://dub1.discourse-cdn.com/flex013/user_avatar/cms-helpdesk.ncas.ac.uk/juan_bilbao/32/458_2.png) [@Juan\_Bilbao](https://cms-helpdesk.ncas.ac.uk/u/Juan_Bilbao)\
**Post date:** [11 August 2026 12:10 UTC](https://cms-helpdesk.ncas.ac.uk/t/walltime-exceeded-time-out-in-atmos-main/2144/6 "2026-08-11T12:10:39Z")

</div>

Hi Isabelle,

Changing the Cylc cycle frequency from `P3M` to `P1M` changes the scheduling frequency but it may not automatically change the UM run length or restart selection.

Can you please check the restart/run-length settings in `u-eb430` compared with `u-dr799`?

what restart dump is selected at the start of the `19790301T0000Z` cycle?

Thanks,  
Juan

---

<div class="post-metadata">

**Author:** ![jonnyhtw](https://dub1.discourse-cdn.com/flex013/user_avatar/cms-helpdesk.ncas.ac.uk/jonnyhtw/32/548_2.png) [@jonnyhtw](https://cms-helpdesk.ncas.ac.uk/u/jonnyhtw)\
**Post date:** [11 August 2026 12:26 UTC](https://cms-helpdesk.ncas.ac.uk/t/walltime-exceeded-time-out-in-atmos-main/2144/7 "2026-08-11T12:26:18Z")

</div>

FYI there appears to be some kind of outage to Monsoon access right now. I’ve sent a Teams message to the maintainers.

Cheers

Jonny

---

<div class="post-metadata">

**Author:** ![isangha](https://avatars.discourse-cdn.com/v4/letter/i/82dd89/32.png) [@isangha](https://cms-helpdesk.ncas.ac.uk/u/isangha)\
**Post date:** [11 August 2026 13:49 UTC](https://cms-helpdesk.ncas.ac.uk/t/walltime-exceeded-time-out-in-atmos-main/2144/8 "2026-08-11T13:49:16Z")

</div>

The dump frequency was 90 days, which I have just changed to 30 days to match the EXPT\_RESUB=‘P1M’ and am rerunning.

The fun-length was changed from 3 months in u-dr799 to 39 years in u-eb430.

I am not sure how to see what restart dump is selected at the start of the 19790301T0000Z cycle, which file would this be in?

Thanks,

Isabelle

---

<div class="post-metadata">

**Author:** ![Juan\_Bilbao](https://dub1.discourse-cdn.com/flex013/user_avatar/cms-helpdesk.ncas.ac.uk/juan_bilbao/32/458_2.png) [@Juan\_Bilbao](https://cms-helpdesk.ncas.ac.uk/u/Juan_Bilbao)\
**Post date:** [11 August 2026 14:06 UTC](https://cms-helpdesk.ncas.ac.uk/t/walltime-exceeded-time-out-in-atmos-main/2144/9 "2026-08-11T14:06:32Z")

</div>

Hi Isabelle,

Thanks for checking this.

Let see if the dump frequency to 30 days to match the monthly resubmission should resolve this.

To check which restart dump is being used, can you check the `atmos_main` `job.out` for the March cycle?

path\_to/19790301T0000Z/atmos\_main/01/job.out

Search for `dump`, `restart`, or `start` messages. Let me know how it goes.

Best regards,

Juan

---

<div class="post-metadata">

**Author:** ![isangha](https://avatars.discourse-cdn.com/v4/letter/i/82dd89/32.png) [@isangha](https://cms-helpdesk.ncas.ac.uk/u/isangha)\
**Post date:** [12 August 2026 08:52 UTC](https://cms-helpdesk.ncas.ac.uk/t/walltime-exceeded-time-out-in-atmos-main/2144/10 "2026-08-12T08:52:51Z")

</div>

Hi Juan,

Changing the dump frequency to 30 days did resolve the issue. Thank you for all your help with this!

---

<div class="post-metadata">

**Author:** ![Juan\_Bilbao](https://dub1.discourse-cdn.com/flex013/user_avatar/cms-helpdesk.ncas.ac.uk/juan_bilbao/32/458_2.png) [@Juan\_Bilbao](https://cms-helpdesk.ncas.ac.uk/u/Juan_Bilbao)\
**Post date:** [12 August 2026 09:17 UTC](https://cms-helpdesk.ncas.ac.uk/t/walltime-exceeded-time-out-in-atmos-main/2144/11 "2026-08-12T09:17:28Z")

</div>

Hi Isabelle,

Good to know that it is now working!

Best regards,  
Juan

---

<div class="post-metadata">

**Author:** ![Juan\_Bilbao](https://dub1.discourse-cdn.com/flex013/user_avatar/cms-helpdesk.ncas.ac.uk/juan_bilbao/32/458_2.png) [@Juan\_Bilbao](https://cms-helpdesk.ncas.ac.uk/u/Juan_Bilbao)\
**Post date:** [12 August 2026 09:17 UTC](https://cms-helpdesk.ncas.ac.uk/t/walltime-exceeded-time-out-in-atmos-main/2144/12 "2026-08-12T09:17:40Z")

</div>


