# BICGstab error 20 years into nudged run

**URL:** <https://cms-helpdesk.ncas.ac.uk/t/bicgstab-error-20-years-into-nudged-run/1519>\
**Category:** Unified Model\
**Tags:** Monsoon2\
**Created:** [10 September 2024 14:51 UTC](https://cms-helpdesk.ncas.ac.uk/t/bicgstab-error-20-years-into-nudged-run/1519 "2024-09-10T14:51:31Z")\
**Posts on this page:** 20\
**Page:** 1

<div class="post-metadata">

**Author:** ![isangha](https://avatars.discourse-cdn.com/v4/letter/i/82dd89/32.png) [@isangha](https://cms-helpdesk.ncas.ac.uk/u/isangha)\
**Post date:** [10 September 2024 14:51 UTC](https://cms-helpdesk.ncas.ac.uk/t/bicgstab-error-20-years-into-nudged-run/1519/1 "2024-09-10T14:51:31Z")

</div>

Hello CMS Helpdesk,

I have been running a nudged run in suite u-df691 with a branch that contains code changes I have added for polar stratospheric cloud formation in the UKCA. The simulation runs fine from 1982 (when it is initialized) to 2000, but fails with a BICGstab error in December 2000 (screenshot attached). Any ideas what could be causing this error so far into a model run?

Thank you!

 ![Screenshot (195)](https://europe1.discourse-cdn.com/flex013/uploads/cms_support/original/1X/17809e8ee34ff2740b2e6bd6dfbef6f63ea60aa8.png)

---

<div class="post-metadata">

**Author:** ![grenville](https://dub1.discourse-cdn.com/flex013/user_avatar/cms-helpdesk.ncas.ac.uk/grenville/32/29_2.png) [@grenville](https://cms-helpdesk.ncas.ac.uk/u/grenville)\
**Post date:** [11 September 2024 07:42 UTC](https://cms-helpdesk.ncas.ac.uk/t/bicgstab-error-20-years-into-nudged-run/1519/2 "2024-09-11T07:42:55Z")

</div>

Isabelle

We generally work around these failures by perturbing the atmosphere start file and rerunning the cycle. See [https://code.metoffice.gov.uk/trac/moci/wiki/tips\_CRgeneral#Restartingifthemodelblowsup](https://code.metoffice.gov.uk/trac/moci/wiki/tips_CRgeneral#Restartingifthemodelblowsup)  
for how to use perturb\_theta.py

Grenville

---

<div class="post-metadata">

**Author:** ![isangha](https://avatars.discourse-cdn.com/v4/letter/i/82dd89/32.png) [@isangha](https://cms-helpdesk.ncas.ac.uk/u/isangha)\
**Post date:** [11 September 2024 09:03 UTC](https://cms-helpdesk.ncas.ac.uk/t/bicgstab-error-20-years-into-nudged-run/1519/3 "2024-09-11T09:03:17Z")

</div>

Hi Grenville,

Thank you. I have looked through this and tried to perturb the theta field, but I am not sure how to use the perturb\_theta.py script on monsoon (ie how to access it from monsoon and how to run the python script).

Thanks!

Best,  
Isabelle

---

<div class="post-metadata">

**Author:** ![mdalvi](https://avatars.discourse-cdn.com/v4/letter/m/ad7895/32.png) [@mdalvi](https://cms-helpdesk.ncas.ac.uk/u/mdalvi)\
**Post date:** [11 September 2024 09:32 UTC](https://cms-helpdesk.ncas.ac.uk/t/bicgstab-error-20-years-into-nudged-run/1519/4 "2024-09-11T09:32:22Z")

</div>

Hi Isabelle,  
The script is not installed by default on Monsoon, so you will have to extract from the repository.  
$ mkdir ~/test (or suitable folder)  
$ cd ~/test/  
$ fcm export fcm:moci.xm\_tr/Utilities/lib/perturb\_theta.py  
$ chmod +x perturb\_theta.py  
$ cd cylc-run/u-df691/share/data/History\_Data/  
(assuming the atmos\_main task starting 01Dec2000 is failing)  
$ mv df691a.da20001201\_00 df691a.da20001201\_00.orig  
$ module load um\_tools  
$ ~/test/perturb\_theta.py df691a.da20001201\_00.orig --output ./df691a.da20001201\_00

Resubmit the failed task.

---

<div class="post-metadata">

**Author:** ![mdalvi](https://avatars.discourse-cdn.com/v4/letter/m/ad7895/32.png) [@mdalvi](https://cms-helpdesk.ncas.ac.uk/u/mdalvi)\
**Post date:** [11 September 2024 10:19 UTC](https://cms-helpdesk.ncas.ac.uk/t/bicgstab-error-20-years-into-nudged-run/1519/6 "2024-09-11T10:19:30Z")

</div>

If the suite has stopped in the meantime, you can restart from the failed point (as long as no configuration settings/ namelists have changed) using:  
$ rose suite-run --restart

---

<div class="post-metadata">

**Author:** ![isangha](https://avatars.discourse-cdn.com/v4/letter/i/82dd89/32.png) [@isangha](https://cms-helpdesk.ncas.ac.uk/u/isangha)\
**Post date:** [12 September 2024 08:41 UTC](https://cms-helpdesk.ncas.ac.uk/t/bicgstab-error-20-years-into-nudged-run/1519/7 "2024-09-12T08:41:46Z")

</div>

Thank you! I have perturbed the dumpfile and resubmitted the task, however now the postproc step seems to be failing (it is stuck in a loop of retrying). The text from job.err file from the postproc task is copied below. Any idea why this would be happening? Thanks in advance!

**update** : the atmos step has failed again for 12/2000 with a BICGstab error

_job.err file :_

[WARN] file:atmospp.nl: skip missing optional source: namelist:archer\_arch  
[WARN] file:pptransfer.nl: skip missing optional source: namelist:archer\_arch  
[WARN] file:atmospp.nl: skip missing optional source: namelist:script\_arch  
[WARN] file:pptransfer.nl: skip missing optional source: namelist:pptransfer  
[WARN] `rose date` requires length=5 input date array - adding 0: [1, 1, 1, 0]  
[WARN] `rose date` requires length=5 input date array - adding 0: [1, 1, 1, 0, 0]  
[WARN] `rose date` requires length=5 input date array - adding 0: [2000, 12, 1, 0]  
[WARN] `rose date` requires length=5 input date array - adding 0: [2000, 12, 1, 0, 0]  
[WARN] `rose date` requires length=5 input date array - adding 0: [2000, 12, 1, 0]  
[WARN] `rose date` requires length=5 input date array - adding 0: [2000, 12, 1, 0, 0]  
[WARN] `rose date` requires length=5 input date array - adding 0: [1, 12, 1, 0]  
[WARN] `rose date` requires length=5 input date array - adding 0: [1, 12, 1, 0, 0]  
[WARN] `rose date` requires length=5 input date array - adding 0: [2000, 12, 1, 0]  
[WARN] `rose date` requires length=5 input date array - adding 0: [2000, 12, 1, 0, 0]  
[WARN] move\_files: Deleted pre-existing file with same name prior to move: /home/d03/isangha/cylc-run/u-df691/work/20001201T0000Z/atmos\_main/df691a.ps2000son.arch  
Traceback (most recent call last):  
File “/home/d03/isangha/cylc-run/u-df691/share/fcm\_make\_pp/build/bin/main\_pp.py”, line 118, in   
main()  
File “/home/d03/isangha/cylc-run/u-df691/share/fcm\_make\_pp/build/bin/main\_pp.py”, line 111, in main  
run\_postproc()  
File “/home/d03/isangha/cylc-run/u-df691/share/fcm\_make\_pp/build/bin/main\_pp.py”, line 82, in run\_postproc  
getattr(model, meth)()  
File “/projects/ukca-cam/isangha/cylc-run/u-df691/share/fcm\_make\_pp/build/bin/timer.py”, line 115, in wrapper  
out = function(\*args, \*\*kw)  
File “/projects/ukca-cam/isangha/cylc-run/u-df691/share/fcm\_make\_pp/build/bin/atmos.py”, line 519, in do\_transform  
for fname in self.diags\_to\_process(finalcycle):  
File “/projects/ukca-cam/isangha/cylc-run/u-df691/share/fcm\_make\_pp/build/bin/timer.py”, line 115, in wrapper  
out = function(\*args, \*\*kw)  
File “/projects/ukca-cam/isangha/cylc-run/u-df691/share/fcm\_make\_pp/build/bin/atmos.py”, line 335, in diags\_to\_process  
logfile=log\_file  
File “/projects/ukca-cam/isangha/cylc-run/u-df691/share/fcm\_make\_pp/build/bin/timer.py”, line 115, in wrapper  
out = function(\*args, \*\*kw)  
File “/projects/ukca-cam/isangha/cylc-run/u-df691/share/fcm\_make\_pp/build/bin/validation.py”, line 178, in verify\_header  
headers, empty\_file = mule\_headers(fname)  
File “/projects/ukca-cam/isangha/cylc-run/u-df691/share/fcm\_make\_pp/build/bin/timer.py”, line 115, in wrapper  
out = function(\*args, \*\*kw)  
File “/projects/ukca-cam/isangha/cylc-run/u-df691/share/fcm\_make\_pp/build/bin/validation.py”, line 283, in mule\_headers  
umfile = mule.UMFile.from\_file(filename, remove\_empty\_lookups=True)  
File “/opt/scitools/environments/production\_legacy/2018\_10\_17/lib/python2.7/site-packages/mule/ **init**.py”, line 1246, in from\_file  
new\_umf.\_read\_file(file\_or\_filepath)  
File “/opt/scitools/environments/production\_legacy/2018\_10\_17/lib/python2.7/site-packages/mule/ **init**.py”, line 1424, in \_read\_file  
FixedLengthHeader.from\_file(source))  
File “/opt/scitools/environments/production\_legacy/2018\_10\_17/lib/python2.7/site-packages/mule/ **init**.py”, line 555, in from\_file  
return super(FixedLengthHeader, cls).from\_file(source, cls.\_NUM\_WORDS)  
File “/opt/scitools/environments/production\_legacy/2018\_10\_17/lib/python2.7/site-packages/mule/ **init**.py”, line 393, in from\_file  
return cls(values)  
File “/opt/scitools/environments/production\_legacy/2018\_10\_17/lib/python2.7/site-packages/mule/ **init**.py”, line 527, in **init**  
raise ValueError(\_msg)  
ValueError: Incorrect size for fixed length header; given 0 words but should be 256.  
[FAIL] main\_pp.py atmos # return-code=1  
2024-09-12T08:29:53Z CRITICAL - failed/EXIT

---

<div class="post-metadata">

**Author:** ![grenville](https://dub1.discourse-cdn.com/flex013/user_avatar/cms-helpdesk.ncas.ac.uk/grenville/32/29_2.png) [@grenville](https://cms-helpdesk.ncas.ac.uk/u/grenville)\
**Post date:** [12 September 2024 12:12 UTC](https://cms-helpdesk.ncas.ac.uk/t/bicgstab-error-20-years-into-nudged-run/1519/8 "2024-09-12T12:12:39Z")

</div>

Hi Isabelle

I think you should have perturbed df691a.da20001201\_00 - that’s the start file that the 20001201T0000Z cylce will use.

Not sure about posproc - let’s get this fixed first.

Grenvile

---

<div class="post-metadata">

**Author:** ![isangha](https://avatars.discourse-cdn.com/v4/letter/i/82dd89/32.png) [@isangha](https://cms-helpdesk.ncas.ac.uk/u/isangha)\
**Post date:** [12 September 2024 12:27 UTC](https://cms-helpdesk.ncas.ac.uk/t/bicgstab-error-20-years-into-nudged-run/1519/9 "2024-09-12T12:27:36Z")

</div>

Hi Grenville,

The dump file I perturbed was the df691a.da20001201\_00 file, however it still results in a BICGstab error when I resubmit the for the 20001201T0000Z cycle.

Best,  
Isabelle

---

<div class="post-metadata">

**Author:** ![mdalvi](https://avatars.discourse-cdn.com/v4/letter/m/ad7895/32.png) [@mdalvi](https://cms-helpdesk.ncas.ac.uk/u/mdalvi)\
**Post date:** [12 September 2024 12:58 UTC](https://cms-helpdesk.ncas.ac.uk/t/bicgstab-error-20-years-into-nudged-run/1519/10 "2024-09-12T12:58:05Z")

</div>

Hi Isabelle,

Just trying to understand the run so far - the postproc task should only launch on successful completion of the atmos\_main task (for that month). Was the postproc task launched manually?  
From looking at your work and share folders it looks like at least part of the simulation was re-run recently (without recon). If the pertrub\_theta does not seem to work for the December restart file, it is likely anamolous values or cause of failure is already ‘baked in’ that dump and it might be worth perturbing an earlier month and re-running the simulation from there.

Mohit

---

<div class="post-metadata">

**Author:** ![isangha](https://avatars.discourse-cdn.com/v4/letter/i/82dd89/32.png) [@isangha](https://cms-helpdesk.ncas.ac.uk/u/isangha)\
**Post date:** [12 September 2024 13:06 UTC](https://cms-helpdesk.ncas.ac.uk/t/bicgstab-error-20-years-into-nudged-run/1519/11 "2024-09-12T13:06:23Z")

</div>

Hi Mohit,

I think the issue with the postproc is that I ran the ‘trigger now’ for the whole cycle 20001201T0000Z rather than just the atmos\_main task and so the postproc was triggered before the atmos\_main finished.

In terms of perturbing an earlier month, I tried perturbing the theta field in November, 2000 restart file and it results in the same BICGstab error in December, 2000. I will try perhaps perturbing a restart file for January, 2000 and see whether that allows the simulation to progress past December, 2000 – although would it take that long for anamolous values to cause a failure?

Thanks!

Best,  
Izzy

---

<div class="post-metadata">

**Author:** ![mdalvi](https://avatars.discourse-cdn.com/v4/letter/m/ad7895/32.png) [@mdalvi](https://cms-helpdesk.ncas.ac.uk/u/mdalvi)\
**Post date:** [12 September 2024 13:22 UTC](https://cms-helpdesk.ncas.ac.uk/t/bicgstab-error-20-years-into-nudged-run/1519/12 "2024-09-12T13:22:49Z")

</div>

If perturbing the values 2 -3 months back still causes a failure at the same point then it is most likely an input happening at/ just before the failure point.  
In case of monthly ancillaries the data for December would have been read from 16th Nov (and interpolated) so would have failed earlier. One other cause could be the Nudging input files, but so far other users have not reported similar problems for this period.

---

<div class="post-metadata">

**Author:** ![grenville](https://dub1.discourse-cdn.com/flex013/user_avatar/cms-helpdesk.ncas.ac.uk/grenville/32/29_2.png) [@grenville](https://cms-helpdesk.ncas.ac.uk/u/grenville)\
**Post date:** [12 September 2024 13:24 UTC](https://cms-helpdesk.ncas.ac.uk/t/bicgstab-error-20-years-into-nudged-run/1519/13 "2024-09-12T13:24:35Z")

</div>

Argh - sorry misread the filename!

---

<div class="post-metadata">

**Author:** ![isangha](https://avatars.discourse-cdn.com/v4/letter/i/82dd89/32.png) [@isangha](https://cms-helpdesk.ncas.ac.uk/u/isangha)\
**Post date:** [12 September 2024 13:30 UTC](https://cms-helpdesk.ncas.ac.uk/t/bicgstab-error-20-years-into-nudged-run/1519/14 "2024-09-12T13:30:42Z")

</div>

Thanks - I will try and perturb September and October and see if the failure still happens.

Just to check that I am re-running these correctly. If I want to perturb the September dumpfile and try rerunning I should 1) perturb the September file using perturb\_theta.py, 2) update the AINITIAL file to be the perturbed September file, 3) Update the model basis time to match the September dumpfile date and turn build and recon off, and 3) run the suite from September.

---

<div class="post-metadata">

**Author:** ![mdalvi](https://avatars.discourse-cdn.com/v4/letter/m/ad7895/32.png) [@mdalvi](https://cms-helpdesk.ncas.ac.uk/u/mdalvi)\
**Post date:** [12 September 2024 13:38 UTC](https://cms-helpdesk.ncas.ac.uk/t/bicgstab-error-20-years-into-nudged-run/1519/15 "2024-09-12T13:38:26Z")

</div>

Yes…  
Note that AINITIAL → Reconf → ASTART. so the Atmos model always reads the ASTART file and not Ainitial. However, the suite may have been set up to link the Ainitial file as Astart directly, in absence of Reconfiguration so need to check.

---

<div class="post-metadata">

**Author:** ![isangha](https://avatars.discourse-cdn.com/v4/letter/i/82dd89/32.png) [@isangha](https://cms-helpdesk.ncas.ac.uk/u/isangha)\
**Post date:** [12 September 2024 13:43 UTC](https://cms-helpdesk.ncas.ac.uk/t/bicgstab-error-20-years-into-nudged-run/1519/16 "2024-09-12T13:43:09Z")

</div>

Ah, I see. The astart is currently : $ROSE\_DATA/${RUNID}$a.astart , but should I set it to /cylc-run/u-df691/share/data/History\_Data/df691a.da20000901\_00 (assuming I want to rerun from September) to ensure that the atmos model is reading in the correct dump file?

Thanks!

---

<div class="post-metadata">

**Author:** ![mdalvi](https://avatars.discourse-cdn.com/v4/letter/m/ad7895/32.png) [@mdalvi](https://cms-helpdesk.ncas.ac.uk/u/mdalvi)\
**Post date:** [12 September 2024 13:45 UTC](https://cms-helpdesk.ncas.ac.uk/t/bicgstab-error-20-years-into-nudged-run/1519/17 "2024-09-12T13:45:23Z")

</div>

Yes, for this run. but be aware that the astart file will be overwritten if Reconfiguration is turned On.

---

<div class="post-metadata">

**Author:** ![isangha](https://avatars.discourse-cdn.com/v4/letter/i/82dd89/32.png) [@isangha](https://cms-helpdesk.ncas.ac.uk/u/isangha)\
**Post date:** [12 September 2024 14:30 UTC](https://cms-helpdesk.ncas.ac.uk/t/bicgstab-error-20-years-into-nudged-run/1519/18 "2024-09-12T14:30:17Z")

</div>

Sorry for the continued errors …  
The run fails at the atmos job when I set ASTART to /cylc-run/u-df691/share/data/History\_Data/df691a.da20001001\_00 with the error copied below. I have tried using ‘/cylc-run/u-df691/share/data/History\_Data/df691a.da20001001\_00’ as well as ‘~/cylc-run/u-df691/share/data/History\_Data/df691a.da20001001\_00’ and both result in the same error.

???  
???!!!???!!!???!!!???!!!???!!! ERROR ???!!!???!!!???!!!???!!!???!!!  
? Error code: 1  
? Error from routine: io:file\_open  
? Error message: Failed to open file /cylc-run/u-df691/share/data/History\_Data/df691a.da20001001\_00  
? Error from processor: 0  
? Error number: 39  
???

---

<div class="post-metadata">

**Author:** ![mdalvi](https://avatars.discourse-cdn.com/v4/letter/m/ad7895/32.png) [@mdalvi](https://cms-helpdesk.ncas.ac.uk/u/mdalvi)\
**Post date:** [12 September 2024 14:41 UTC](https://cms-helpdesk.ncas.ac.uk/t/bicgstab-error-20-years-into-nudged-run/1519/19 "2024-09-12T14:41:12Z")

</div>

Hi,  
You will have to specify the full path starting from /projects/ukca-cam/ (i.e $DATADIR)/cylc-run/suite-id/. That is usually where the share/data folders are installed.

---

<div class="post-metadata">

**Author:** ![isangha](https://avatars.discourse-cdn.com/v4/letter/i/82dd89/32.png) [@isangha](https://cms-helpdesk.ncas.ac.uk/u/isangha)\
**Post date:** [13 September 2024 07:28 UTC](https://cms-helpdesk.ncas.ac.uk/t/bicgstab-error-20-years-into-nudged-run/1519/20 "2024-09-13T07:28:14Z")

</div>

Hi Mohit,

Thank you, I think that the astart file can now be found however it is failing with a ‘cylc: unbound variable error’ related to the astart file.

 ![Screenshot (196)](https://europe1.discourse-cdn.com/flex013/uploads/cms_support/original/1X/65e50c4ee918f9a9c3376296f42fe100a10109ce.png)

---

<div class="post-metadata">

**Author:** ![mdalvi](https://avatars.discourse-cdn.com/v4/letter/m/ad7895/32.png) [@mdalvi](https://cms-helpdesk.ncas.ac.uk/u/mdalvi)\
**Post date:** [13 September 2024 09:38 UTC](https://cms-helpdesk.ncas.ac.uk/t/bicgstab-error-20-years-into-nudged-run/1519/21 "2024-09-13T09:38:43Z")

</div>

Hi,  
My earlier reply only contained an indication of what the full path should be (as I was not sure of what your $DATADIR is !).  
The following setting should work:  
astart=‘/projects/ukca-cam/isangha/cylc-run/u-df691/share/data/History\_Data/df691a.da20001001\_00’

[Next page](https://cms-helpdesk.ncas.ac.uk/t/bicgstab-error-20-years-into-nudged-run/1519.md?page=2)
