# ARCHER2 Rename issue

**URL:** <https://cms-helpdesk.ncas.ac.uk/t/archer2-rename-issue/248>\
**Category:** Unified Model\
**Tags:** ARCHER2\
**Created:** [6 October 2021 11:45 UTC](https://cms-helpdesk.ncas.ac.uk/t/archer2-rename-issue/248 "2021-10-06T11:45:57Z")\
**Posts on this page:** 10\
**Page:** 1

<div class="post-metadata">

**Author:** ![dgaleareading](https://dub1.discourse-cdn.com/flex013/user_avatar/cms-helpdesk.ncas.ac.uk/dgaleareading/32/30_2.png) [@dgaleareading](https://cms-helpdesk.ncas.ac.uk/u/dgaleareading)\
**Post date:** [6 October 2021 11:45 UTC](https://cms-helpdesk.ncas.ac.uk/t/archer2-rename-issue/248/1 "2021-10-06T11:45:57Z")

</div>

Hi,

I have been running the UM using the old archer2 name (`login.archer.ac.uk`) but I think today’s change to `login-4c.archer.ac.uk` may be causing an issue. I have attached a screenshot of the cylc-gui for my run showing a supposedly running pptransfer task, however this has finished on ARCHER2. When I re-poll the tasks, its status is not updated. Also, when trying to get a job status or error file, I get the error seen. I have now changed the login address in both site/archer2.rc and in the ssh config file. However, when trying to run `suite-run reload`, I get that the suite still has running tasks. How do you think I should proceed?

Regards.  
Daniel

 ![Screenshot from 2021-10-06 12-43-52](https://europe1.discourse-cdn.com/flex013/uploads/cms_support/original/1X/725794f8763295692e5dbbf799abd790ea5ed52f.png)

---

<div class="post-metadata">

**Author:** ![RosalynHatcher](https://dub1.discourse-cdn.com/flex013/user_avatar/cms-helpdesk.ncas.ac.uk/rosalynhatcher/32/11_2.png) [@RosalynHatcher](https://cms-helpdesk.ncas.ac.uk/u/RosalynHatcher)\
**Post date:** [6 October 2021 12:10 UTC](https://cms-helpdesk.ncas.ac.uk/t/archer2-rename-issue/248/2 "2021-10-06T12:10:17Z")

</div>

Hi Daniel,

I would suggest manually checking the `job.status` file for the `pptransfer` task on ARCHER2. If the status is succeeded then in the cylc GUI change the status of `pptransfer` to succeeded. The next tasks should then start up and use `login-4c` from then on. Are you sure it was `rose suite-run --reload` you ran as that shouldn’t mind if tasks are already running. If it doesn’t work I would suggest stopping and then restarting the suite with `rose suite-run --restart`.

Have you also updated your `~/.ssh/config` file on PUMA to change the `login.archer2.ac.uk` to `login-4c.archer2.ac.uk`?

Cheers,  
Ros.

---

<div class="post-metadata">

**Author:** ![dgaleareading](https://dub1.discourse-cdn.com/flex013/user_avatar/cms-helpdesk.ncas.ac.uk/dgaleareading/32/30_2.png) [@dgaleareading](https://cms-helpdesk.ncas.ac.uk/u/dgaleareading)\
**Post date:** [6 October 2021 13:04 UTC](https://cms-helpdesk.ncas.ac.uk/t/archer2-rename-issue/248/3 "2021-10-06T13:04:37Z")

</div>

Hi,

stopping the suite and running `rose suite-run --restart` seems to have done the trick. Thanks.

Regards,  
Daniel

---

<div class="post-metadata">

**Author:** ![dgaleareading](https://dub1.discourse-cdn.com/flex013/user_avatar/cms-helpdesk.ncas.ac.uk/dgaleareading/32/30_2.png) [@dgaleareading](https://cms-helpdesk.ncas.ac.uk/u/dgaleareading)\
**Post date:** [7 October 2021 09:37 UTC](https://cms-helpdesk.ncas.ac.uk/t/archer2-rename-issue/248/4 "2021-10-07T09:37:35Z")

</div>

Hi,

It seems that my pptransfer task is stuck on “retrying” due to the following error:

```auto
connection failed
connection failed
connection failed
connection failed
connection failed
connection failed
connection failed
WARNING: MESSAGE SEND FAILED
CPU time limit exceeded
Received signal XCPU
cylc (scheduler - 2021-10-07T03:23:37Z): CRITICAL Task job script received signal XCPU at 2021-10-07T03:23:37Z
cylc (scheduler - 2021-10-07T03:23:37Z): CRITICAL failed at 2021-10-07T03:23:37Z
connection failed
connection failed
connection failed
connection failed
connection failed
connection failed
connection failed
WARNING: MESSAGE SEND FAILED

```

Could this still be linked to the ARCHER2 name change? I have tried ssh’ing to JASMIN from ARCHER2 via `ssh hpxfer1.jasmin.ac.uk` and that worked fine.

Regards,  
Daniel

---

<div class="post-metadata">

**Author:** ![RosalynHatcher](https://dub1.discourse-cdn.com/flex013/user_avatar/cms-helpdesk.ncas.ac.uk/rosalynhatcher/32/11_2.png) [@RosalynHatcher](https://cms-helpdesk.ncas.ac.uk/u/RosalynHatcher)\
**Post date:** [7 October 2021 10:50 UTC](https://cms-helpdesk.ncas.ac.uk/t/archer2-rename-issue/248/5 "2021-10-07T10:50:31Z")

</div>

Hi Daniel,

Yes, I’ve just tweaked the cylc configuration file again on PUMA so cylc should now try polling rather than trying to communicate back from ARCHER2 which it can’t do. Please try retriggering the task. If it still does the same I would suggest stopping and restarting the suite.

Cheers,  
Ros.

---

<div class="post-metadata">

**Author:** ![dgaleareading](https://dub1.discourse-cdn.com/flex013/user_avatar/cms-helpdesk.ncas.ac.uk/dgaleareading/32/30_2.png) [@dgaleareading](https://cms-helpdesk.ncas.ac.uk/u/dgaleareading)\
**Post date:** [7 October 2021 16:23 UTC](https://cms-helpdesk.ncas.ac.uk/t/archer2-rename-issue/248/6 "2021-10-07T16:23:57Z")

</div>

Hi,

I’ve stopped and restarted the suite, but the same error still crops up.

Regards,  
Daniel

---

<div class="post-metadata">

**Author:** ![RosalynHatcher](https://dub1.discourse-cdn.com/flex013/user_avatar/cms-helpdesk.ncas.ac.uk/rosalynhatcher/32/11_2.png) [@RosalynHatcher](https://cms-helpdesk.ncas.ac.uk/u/RosalynHatcher)\
**Post date:** [8 October 2021 07:45 UTC](https://cms-helpdesk.ncas.ac.uk/t/archer2-rename-issue/248/7 "2021-10-08T07:45:24Z")

</div>

HI Daniel,

So it’s fixed the communication back problem. The suite is now polling and we’ve lost the `connection failed` errors. The problem now is that it’s just not completing in the allocated time. I was watching the checksumming last night and it took about 2 hours and then only managed to transfer about 80gb to JASMIN in the next 2 hours. I don’t know if this is ARCHER2 load on the login nodes, filesystem issues or connection to JASMIN. I’m going to take a look at another run to see if that’s seeing any slowdown.

Cheers,  
Ros.

---

<div class="post-metadata">

**Author:** ![RosalynHatcher](https://dub1.discourse-cdn.com/flex013/user_avatar/cms-helpdesk.ncas.ac.uk/rosalynhatcher/32/11_2.png) [@RosalynHatcher](https://cms-helpdesk.ncas.ac.uk/u/RosalynHatcher)\
**Post date:** [8 October 2021 09:03 UTC](https://cms-helpdesk.ncas.ac.uk/t/archer2-rename-issue/248/8 "2021-10-08T09:03:59Z")

</div>

Hi Daniel,

Just been looking at your suite further. `19910401T0000Z` cycle still has only transferred 76Gb which is the same as when I looked at it at late last night, so it looks like the rsync has got stuck for some reason. I’m currently running a 300Gb transfer and it’s already done 100Gb so I don’t think it’s the connection to JASMIN.

Can you please try killing the `19910401T0000Z/pptransfer` task that is currently running. Then on JASMIN move the  
`/gws/nopw/j04/hiresgw/dg/archer_transfers/u-cg647/19910401T0000Z` directory out of the way and then retrigger the `pptransfer` task and see if a fresh run of the task clears it.

Cheers,  
Ros.

---

<div class="post-metadata">

**Author:** ![dgaleareading](https://dub1.discourse-cdn.com/flex013/user_avatar/cms-helpdesk.ncas.ac.uk/dgaleareading/32/30_2.png) [@dgaleareading](https://cms-helpdesk.ncas.ac.uk/u/dgaleareading)\
**Post date:** [8 October 2021 09:22 UTC](https://cms-helpdesk.ncas.ac.uk/t/archer2-rename-issue/248/9 "2021-10-08T09:22:19Z")

</div>

Hi,

I now think that there is a lack of space in my gws on JASMIN. I’ve removed some of the dumps from the earlier cycles and the rsync task on ARCHER2 has started going again, with some new files on JASMIN.

Regards,  
Daniel

---

<div class="post-metadata">

**Author:** ![system](https://europe1.discourse-cdn.com/flex013/uploads/cms_support/original/1X/1fd2411499ffcbc299fe756cd5cdf26e44956558.png) [@system](https://cms-helpdesk.ncas.ac.uk/u/system)\
**Post date:** [10 October 2021 09:22 UTC](https://cms-helpdesk.ncas.ac.uk/t/archer2-rename-issue/248/10 "2021-10-10T09:22:46Z")

</div>

This topic was automatically closed 2 days after the last reply. New replies are no longer allowed.
