# Fcm\_make2\_um failing

**URL:** <https://cms-helpdesk.ncas.ac.uk/t/fcm-make2-um-failing/1395>\
**Category:** Unified Model\
**Tags:** ARCHER2\
**Created:** [26 April 2024 14:58 UTC](https://cms-helpdesk.ncas.ac.uk/t/fcm-make2-um-failing/1395 "2024-04-26T14:58:48Z")\
**Posts on this page:** 18\
**Page:** 1

<div class="post-metadata">

**Author:** ![c.j.r.williams](https://dub1.discourse-cdn.com/flex013/user_avatar/cms-helpdesk.ncas.ac.uk/c.j.r.williams/32/591_2.png) [@c.j.r.williams](https://cms-helpdesk.ncas.ac.uk/u/c.j.r.williams)\
**Post date:** [26 April 2024 14:58 UTC](https://cms-helpdesk.ncas.ac.uk/t/fcm-make2-um-failing/1395/1 "2024-04-26T14:58:49Z")

</div>

Hi,

Having managed to successfully submit a suite now (many thanks to previous response), and identifying/correcting a couple of failures, it is now failing at the fcm\_make2\_um stage and I don’t know why. There is nothing obvious in any of the output files, or at least nothing that is similar to any of the other comments here. What have I done wrong?

My suite is u-df570.

Thanks,

Charlie

---

<div class="post-metadata">

**Author:** ![grenville](https://dub1.discourse-cdn.com/flex013/user_avatar/cms-helpdesk.ncas.ac.uk/grenville/32/29_2.png) [@grenville](https://cms-helpdesk.ncas.ac.uk/u/grenville)\
**Post date:** [26 April 2024 15:33 UTC](https://cms-helpdesk.ncas.ac.uk/t/fcm-make2-um-failing/1395/2 "2024-04-26T15:33:36Z")

</div>

Hi Charlie

This is a bit odd. The fcm\_make2\_um task appeared to succeeded on its first try but left a lock file behind. The second try failed because the lock file is present.

The err file says:  
`[FAIL] /work/n02/n02/cjrw09/cylc-run/u-df570/share/fcm_make_um/fcm-make2.lock: lock exists at the destination`

remove the lock file (it is a directory) & retrigger

Grenville

---

<div class="post-metadata">

**Author:** ![c.j.r.williams](https://dub1.discourse-cdn.com/flex013/user_avatar/cms-helpdesk.ncas.ac.uk/c.j.r.williams/32/591_2.png) [@c.j.r.williams](https://cms-helpdesk.ncas.ac.uk/u/c.j.r.williams)\
**Post date:** [3 May 2024 14:09 UTC](https://cms-helpdesk.ncas.ac.uk/t/fcm-make2-um-failing/1395/3 "2024-05-03T14:09:10Z")

</div>

Thanks Grenville, and sorry for the delay in getting back to you. I don’t know why I always get to this on a Friday afternoon - probably because it takes me that long to work through all my other jobs during the week, all of which seem to require immediate responses!

Anyway, I don’t understand this, because if I look in that directory on Archer2, there is no obvious lock file:

cjrw09@ln01:~\> ls /work/n02/n02/cjrw09/cylc-run/u-df570/share/fcm\_make\_um/

build-atmos extract fcm-make2.cfg fcm-make2.log preprocess-atmos  
build-recon fcm-make2-as-parsed.cfg fcm-make2.cfg.orig fcm-make2-on-success.cfg preprocess-recon

I certainly haven’t removed anything, since we spoke.

But I have now tried running again, and this time it got passed that stage no problem. So it seems to have now built okay, and the reconfiguration is queueing. When (because knowing my luck, it won’t be if) this fails as well, I will open a new ticket if I can’t solve the error myself.

Charlie

---

<div class="post-metadata">

**Author:** ![c.j.r.williams](https://dub1.discourse-cdn.com/flex013/user_avatar/cms-helpdesk.ncas.ac.uk/c.j.r.williams/32/591_2.png) [@c.j.r.williams](https://cms-helpdesk.ncas.ac.uk/u/c.j.r.williams)\
**Post date:** [3 May 2024 15:19 UTC](https://cms-helpdesk.ncas.ac.uk/t/fcm-make2-um-failing/1395/4 "2024-05-03T15:19:14Z")

</div>

Hi again,

As I feared, it failed at the coupled stage. But at least this error is potentially easy to fix, if indeed this is the error and not a red herring:

BUFFOUT: Write Failed: Disk quota exceeded

FLUSH\_UNIT\_BUFFER: Error Flushing Buffered Data on PE 0  
FLUSH\_UNIT\_BUFFER: Status is 1.0  
FLUSH\_UNIT\_BUFFER: Length Requested was 524288  
FLUSH\_UNIT\_BUFFER: Length written was 1024

It didn’t even get to the postprocessing or pptransfer stage, so this isn’t a JASMIN problem I guess? Do I not have enough temporary space on ARCHER2?

Charlie

---

<div class="post-metadata">

**Author:** ![grenville](https://dub1.discourse-cdn.com/flex013/user_avatar/cms-helpdesk.ncas.ac.uk/grenville/32/29_2.png) [@grenville](https://cms-helpdesk.ncas.ac.uk/u/grenville)\
**Post date:** [9 May 2024 10:32 UTC](https://cms-helpdesk.ncas.ac.uk/t/fcm-make2-um-failing/1395/5 "2024-05-09T10:32:08Z")

</div>

Hi Charlie

I increased your quota to 1TB - there’s more, but as usual n02 is filling up.

Grenville

---

<div class="post-metadata">

**Author:** ![c.j.r.williams](https://dub1.discourse-cdn.com/flex013/user_avatar/cms-helpdesk.ncas.ac.uk/c.j.r.williams/32/591_2.png) [@c.j.r.williams](https://cms-helpdesk.ncas.ac.uk/u/c.j.r.williams)\
**Post date:** [10 May 2024 09:18 UTC](https://cms-helpdesk.ncas.ac.uk/t/fcm-make2-um-failing/1395/6 "2024-05-10T09:18:38Z")

</div>

Hi,

Thanks very much indeed. However, it crashed again last night, after queueing for ages, with what I think is the following error:

slurmstepd: error: \*\*\* STEP 6532490.0+0 ON nid003986 CANCELLED AT 2024-05-10T09:33:38 DUE TO TIME LIMIT \*\*\*  
slurmstepd: error: \*\*\* STEP 6532490.0+2 ON nid004033 CANCELLED AT 2024-05-10T09:33:38 DUE TO TIME LIMIT \*\*\*  
slurmstepd: error: \*\*\* STEP 6532490.0+1 ON nid004024 CANCELLED AT 2024-05-10T09:33:38 DUE TO TIME LIMIT \*\*\*  
slurmstepd: error: \*\*\* JOB 6532490 ON nid003986 CANCELLED AT 2024-05-10T09:33:38 DUE TO TIME LIMIT \*\*\*

I don’t think it had actually written anything out.

Charlie

---

<div class="post-metadata">

**Author:** ![RosalynHatcher](https://dub1.discourse-cdn.com/flex013/user_avatar/cms-helpdesk.ncas.ac.uk/rosalynhatcher/32/11_2.png) [@RosalynHatcher](https://cms-helpdesk.ncas.ac.uk/u/RosalynHatcher)\
**Post date:** [10 May 2024 10:51 UTC](https://cms-helpdesk.ncas.ac.uk/t/fcm-make2-um-failing/1395/7 "2024-05-10T10:51:50Z")

</div>

Hi Charlie,

You need to increase the time limit. The model has got as far as 13/11/1850 in 6hours so just needs a bit more time to complete the year you have requested. Alternatively you could reduce the cycling period.

I’ve found this information in `/home/n02/n02/cjrw09/cylc-run/u-df570/work/18500101T0000Z/coupled/pe_output/df570.fort6.pe000`

All the model output is under` /home/n02/n02/cjrw09/cylc-run/u-df570/share/data/History_Data`

Regards,  
Ros.

---

<div class="post-metadata">

**Author:** ![c.j.r.williams](https://dub1.discourse-cdn.com/flex013/user_avatar/cms-helpdesk.ncas.ac.uk/c.j.r.williams/32/591_2.png) [@c.j.r.williams](https://cms-helpdesk.ncas.ac.uk/u/c.j.r.williams)\
**Post date:** [10 May 2024 11:07 UTC](https://cms-helpdesk.ncas.ac.uk/t/fcm-make2-um-failing/1395/8 "2024-05-10T11:07:52Z")

</div>

Hi Ros,

Thanks very much indeed. That’s very odd, though, because I modelled my suite on the PI from Seb S, so presumed it would run at the same speed. Given that it almost finished my test year, how much you think I should increase the time to? 8 hours? 12 hours? I’m a bit unsure how the queueing system works now, do we still have different queues with different priorities?

Thanks,

Charlie

---

<div class="post-metadata">

**Author:** ![RosalynHatcher](https://dub1.discourse-cdn.com/flex013/user_avatar/cms-helpdesk.ncas.ac.uk/rosalynhatcher/32/11_2.png) [@RosalynHatcher](https://cms-helpdesk.ncas.ac.uk/u/RosalynHatcher)\
**Post date:** [10 May 2024 11:55 UTC](https://cms-helpdesk.ncas.ac.uk/t/fcm-make2-um-failing/1395/9 "2024-05-10T11:55:25Z")

</div>

Hi Charlie,

Figuring out the best wallclock is trial and error I’m afraid. Try 8 or 9 hours. You don’t want to cut the time too finely to allow for system jitter. Yet, you don’t want to allow too much excess time as it may impact on the time spent queueing.

The main compute queue is the standard queue (QoS).  
Detail on all the ARCHER2 queues is available here: [Running jobs - ARCHER2 User Documentation](https://docs.archer2.ac.uk/user-guide/scheduler/#resource-limits)

Cheers,  
Ros

---

<div class="post-metadata">

**Author:** ![c.j.r.williams](https://dub1.discourse-cdn.com/flex013/user_avatar/cms-helpdesk.ncas.ac.uk/c.j.r.williams/32/591_2.png) [@c.j.r.williams](https://cms-helpdesk.ncas.ac.uk/u/c.j.r.williams)\
**Post date:** [10 May 2024 12:10 UTC](https://cms-helpdesk.ncas.ac.uk/t/fcm-make2-um-failing/1395/10 "2024-05-10T12:10:00Z")

</div>

Okay, thanks, all understood. I am just running now with 8 hours, so fingers crossed.

Charlie

---

<div class="post-metadata">

**Author:** ![c.j.r.williams](https://dub1.discourse-cdn.com/flex013/user_avatar/cms-helpdesk.ncas.ac.uk/c.j.r.williams/32/591_2.png) [@c.j.r.williams](https://cms-helpdesk.ncas.ac.uk/u/c.j.r.williams)\
**Post date:** [11 May 2024 09:51 UTC](https://cms-helpdesk.ncas.ac.uk/t/fcm-make2-um-failing/1395/11 "2024-05-11T09:51:42Z")

</div>

Sorry Ros, it has failed again at the coupled stage, despite (as far as I can see) completing the year correctly i.e. writing out everything it should. I can’t see any obvious error this time, so what has happened?

Charlie

---

<div class="post-metadata">

**Author:** ![RosalynHatcher](https://dub1.discourse-cdn.com/flex013/user_avatar/cms-helpdesk.ncas.ac.uk/rosalynhatcher/32/11_2.png) [@RosalynHatcher](https://cms-helpdesk.ncas.ac.uk/u/RosalynHatcher)\
**Post date:** [11 May 2024 20:21 UTC](https://cms-helpdesk.ncas.ac.uk/t/fcm-make2-um-failing/1395/12 "2024-05-11T20:21:42Z")

</div>

Hi Charlie,

You hit your /work disk quota - see job.err file.

I’ve just given you a bit more space, but n02 is currently creaking at the seams at the moment. Please resubmit and it should be fine this time.

Cheers  
Ros.

---

<div class="post-metadata">

**Author:** ![c.j.r.williams](https://dub1.discourse-cdn.com/flex013/user_avatar/cms-helpdesk.ncas.ac.uk/c.j.r.williams/32/591_2.png) [@c.j.r.williams](https://cms-helpdesk.ncas.ac.uk/u/c.j.r.williams)\
**Post date:** [13 May 2024 11:01 UTC](https://cms-helpdesk.ncas.ac.uk/t/fcm-make2-um-failing/1395/13 "2024-05-13T11:01:27Z")

</div>

Thank you very much indeed. That’s odd, though, because Grenville only gave me some extra space a couple of days ago. Or was that for /home?

What is a reasonable amount of room to have on /work? At the moment, apart from my current run on Friday (at /work/n02/n02/cjrw09/cylc-run/u-df570) - which itself is a surprisingly large 937 G - the only other data I have on there are 65 G (which is everything I copied over from NEXCS, when we transitioned to ARCHER2). Is that amount okay?

Charlie

---

<div class="post-metadata">

**Author:** ![c.j.r.williams](https://dub1.discourse-cdn.com/flex013/user_avatar/cms-helpdesk.ncas.ac.uk/c.j.r.williams/32/591_2.png) [@c.j.r.williams](https://cms-helpdesk.ncas.ac.uk/u/c.j.r.williams)\
**Post date:** [14 May 2024 11:19 UTC](https://cms-helpdesk.ncas.ac.uk/t/fcm-make2-um-failing/1395/14 "2024-05-14T11:19:56Z")

</div>

Hi again,

Sorry, but although it has now finished the coupled stage, it has now crashed at the postproc\_atmos and postproc\_nemo stages. I can see another time limit error, is this the problem?

Many thanks,

Charlie

PS. Let me know if I should start a new ticket with this, as it is technically a new problem?

---

<div class="post-metadata">

**Author:** ![RosalynHatcher](https://dub1.discourse-cdn.com/flex013/user_avatar/cms-helpdesk.ncas.ac.uk/rosalynhatcher/32/11_2.png) [@RosalynHatcher](https://cms-helpdesk.ncas.ac.uk/u/RosalynHatcher)\
**Post date:** [14 May 2024 11:50 UTC](https://cms-helpdesk.ncas.ac.uk/t/fcm-make2-um-failing/1395/15 "2024-05-14T11:50:36Z")

</div>

Hi Charlie,

Yes postproc\_atmos & postproc\_nemo have run out of time. 2 hours won’t be enough time to processing a year’s worth of data. You’ll need to increase the timelimit.

Change the value of “`execution time limit`” in `site/archer2.rc` within section `[[POSTPROC_RESOURCE]]`

Regards,  
Ros.

P.S. If you still have problems, yes please do start a new ticket.

---

<div class="post-metadata">

**Author:** ![c.j.r.williams](https://dub1.discourse-cdn.com/flex013/user_avatar/cms-helpdesk.ncas.ac.uk/c.j.r.williams/32/591_2.png) [@c.j.r.williams](https://cms-helpdesk.ncas.ac.uk/u/c.j.r.williams)\
**Post date:** [14 May 2024 12:05 UTC](https://cms-helpdesk.ncas.ac.uk/t/fcm-make2-um-failing/1395/16 "2024-05-14T12:05:11Z")

</div>

Thanks very much, I will do that now. Do you have any feeling for how long it needs to do a year? Once I have increased the time, can I just retrigger it, or do I need to start all over again?

Charlie

---

<div class="post-metadata">

**Author:** ![RosalynHatcher](https://dub1.discourse-cdn.com/flex013/user_avatar/cms-helpdesk.ncas.ac.uk/rosalynhatcher/32/11_2.png) [@RosalynHatcher](https://cms-helpdesk.ncas.ac.uk/u/RosalynHatcher)\
**Post date:** [14 May 2024 12:21 UTC](https://cms-helpdesk.ncas.ac.uk/t/fcm-make2-um-failing/1395/17 "2024-05-14T12:21:22Z")

</div>

Hi Charlie,

The files that have so far been processed are in:  
`/work/n02/n02/cjrw09/archive/u-df570/18500101T0000Z`.

So depending on what streams you have active you might be able to get a rough idea of how much extra time, otherwise again it’s just trial and error.

Change the timelimit, reload the suite and then retrigger the task.

Cheers,  
Ros.

---

<div class="post-metadata">

**Author:** ![system](https://europe1.discourse-cdn.com/flex013/uploads/cms_support/original/1X/1fd2411499ffcbc299fe756cd5cdf26e44956558.png) [@system](https://cms-helpdesk.ncas.ac.uk/u/system)\
**Post date:** [13 June 2024 12:21 UTC](https://cms-helpdesk.ncas.ac.uk/t/fcm-make2-um-failing/1395/18 "2024-06-13T12:21:23Z")

</div>

This topic was automatically closed 30 days after the last reply. New replies are no longer allowed.
