# Remote\_setup failing

**URL:** <https://cms-helpdesk.ncas.ac.uk/t/remote-setup-failing/1764>\
**Category:** Unified Model\
**Tags:** PUMA, ARCHER2\
**Created:** [1 August 2025 15:54 UTC](https://cms-helpdesk.ncas.ac.uk/t/remote-setup-failing/1764 "2025-08-01T15:54:42Z")\
**Posts on this page:** 14\
**Page:** 1

<div class="post-metadata">

**Author:** ![SimonTett](https://avatars.discourse-cdn.com/v4/letter/s/3da27b/32.png) [@SimonTett](https://cms-helpdesk.ncas.ac.uk/u/SimonTett)\
**Post date:** [1 August 2025 15:54 UTC](https://cms-helpdesk.ncas.ac.uk/t/remote-setup-failing/1764/1 "2025-08-01T15:54:42Z")

</div>

I am running a cylc 8 suite – u-dr783 and I am getting a failure from remote\_setup.

Looking in the workflow log I have a bunch of errors like:

> 2025-08-01T15:45:42Z WARNING - platform: ln02 - Could not connect to ln02.  
> \* ln02 has been added to the list of unreachable hosts  
> \* remote-init will retry if another host is available.

which suggests, to me, cylc is having trouble connecting to ln02. If I do ssh ln02 from puma I am asked for my archer2 ssh passphrase. If that is happening for cylc then that would explain the failure.

What do I do?

Simon

---

<div class="post-metadata">

**Author:** ![RosalynHatcher](https://dub1.discourse-cdn.com/flex013/user_avatar/cms-helpdesk.ncas.ac.uk/rosalynhatcher/32/11_2.png) [@RosalynHatcher](https://cms-helpdesk.ncas.ac.uk/u/RosalynHatcher)\
**Post date:** [2 August 2025 16:27 UTC](https://cms-helpdesk.ncas.ac.uk/t/remote-setup-failing/1764/2 "2025-08-02T16:27:46Z")

</div>

Hi Simon,

This suggests that your `ssh-agent` on PUMA2 has died.

See here for how to fix it: [11. Appendix B: SSH FAQs — NCAS Unified Model Introduction](https://ncas-cms.github.io/um-training/ssh-tasks.html#restarting-agent)

Cheers,  
Ros

---

<div class="post-metadata">

**Author:** ![SimonTett](https://avatars.discourse-cdn.com/v4/letter/s/3da27b/32.png) [@SimonTett](https://cms-helpdesk.ncas.ac.uk/u/SimonTett)\
**Post date:** [4 August 2025 08:33 UTC](https://cms-helpdesk.ncas.ac.uk/t/remote-setup-failing/1764/3 "2025-08-04T08:33:13Z")

</div>

Did that and ssh ln01 works as expected. However, looks my cylc jobs u-dr157/run7 is behaving strangely. u-dr157 thinks it is running atmos\_main since 14:53 on the 1st of August. squeue –me shows no atmos\_main running… And re-polling doesn’t change anything.

My other job, u-dr783, had failures from remote\_setup. I got u-dr783 to start by triggering remote\_setup – which has ran. I then got a submit failure from install\_ancil so I triggered that again which failed again. Error is:

ssh -oBatchMode=yes -oConnectTimeout=8 -oStrictHostKeyChecking=no ln01 env CYLC\_VERSION=8.4.4 CYLC\_ENV\_NAME=cylc-8.4.4-1 bash --login -c ‘’“‘”‘exec “$0” “$@”’“’”‘’ cylc jobs-submit --utc-mode --remote-mode --clean-env --path=/bin --path=/usr/bin --path=/usr/local/bin --path=/sbin --path=/usr/sbin --path=/usr/local/sbin – ‘$HOME/cylc-run/u-dr783/log/job’ 19790101T0000Z/install\_ancil/01  
[jobs-submit ret\_code] 1  
[jobs-submit out] 2025-08-04T08:27:45Z|19790101T0000Z/install\_ancil/01|1|None

If I, interactively, do ssh -oBatchMode=yes -oConnectTimeout=8 -oStrictHostKeyChecking=no ln01 ls

I get an error:

Warning: Permanently added ‘ln01,10.252.1.65’ (ECDSA) to the list of known hosts.  
tetts@ln01: Permission denied (keyboard-interactive).

Simon

---

<div class="post-metadata">

**Author:** ![RosalynHatcher](https://dub1.discourse-cdn.com/flex013/user_avatar/cms-helpdesk.ncas.ac.uk/rosalynhatcher/32/11_2.png) [@RosalynHatcher](https://cms-helpdesk.ncas.ac.uk/u/RosalynHatcher)\
**Post date:** [5 August 2025 07:46 UTC](https://cms-helpdesk.ncas.ac.uk/t/remote-setup-failing/1764/4 "2025-08-05T07:46:42Z")

</div>

The problem with `install_ancil` submission is this:

```auto
2025-08-04T08:28:24Z [STDERR] sbatch: error: AssocMaxCpuMinutesPerJobLimit
2025-08-04T08:28:24Z [STDERR] sbatch: error: Batch job submission failed: Job violates accounting/QOS policy (job submit limit, user's size and/or time limits)

```

n02-TERRAFIRMA has used all its current budget. I’ll sort it out when I get to the office shortly.

Cheers,  
Ros

---

<div class="post-metadata">

**Author:** ![SimonTett](https://avatars.discourse-cdn.com/v4/letter/s/3da27b/32.png) [@SimonTett](https://cms-helpdesk.ncas.ac.uk/u/SimonTett)\
**Post date:** [5 August 2025 07:54 UTC](https://cms-helpdesk.ncas.ac.uk/t/remote-setup-failing/1764/5 "2025-08-05T07:54:59Z")

</div>

Hi Roz,

thanks a lot! So many ways in which things can go wrong. I don’t think it is me burning the budget!

Simon

---

<div class="post-metadata">

**Author:** ![RosalynHatcher](https://dub1.discourse-cdn.com/flex013/user_avatar/cms-helpdesk.ncas.ac.uk/rosalynhatcher/32/11_2.png) [@RosalynHatcher](https://cms-helpdesk.ncas.ac.uk/u/RosalynHatcher)\
**Post date:** [5 August 2025 08:29 UTC](https://cms-helpdesk.ncas.ac.uk/t/remote-setup-failing/1764/6 "2025-08-05T08:29:14Z")

</div>

All topped up. Indeed, you’re not the biggest user. 🙂

---

<div class="post-metadata">

**Author:** ![SimonTett](https://avatars.discourse-cdn.com/v4/letter/s/3da27b/32.png) [@SimonTett](https://cms-helpdesk.ncas.ac.uk/u/SimonTett)\
**Post date:** [5 August 2025 08:44 UTC](https://cms-helpdesk.ncas.ac.uk/t/remote-setup-failing/1764/7 "2025-08-05T08:44:03Z")

</div>

Great. And how do I get u-dr157/run7 going again? Should I just trigger the next atmos\_main task?

From the log 19801001T000Z/atmos\_main has ran. The .err file has a message at the end:

> /work/n02/n02/tetts/cylc-run/u-dr157/run7/bin/save\_wallclock.sh: /work/n02/n02/  
> tetts/cylc-run/u-dr157/run7/bin/iteration\_bins.py: /usr/bin/python: bad interpreter: No such file or directory

I could fix that by modifying iteration\_bins.py but I think that this does not matter… It is in the logs for other atmos\_main and they worked.

Simon

---

<div class="post-metadata">

**Author:** ![SimonTett](https://avatars.discourse-cdn.com/v4/letter/s/3da27b/32.png) [@SimonTett](https://cms-helpdesk.ncas.ac.uk/u/SimonTett)\
**Post date:** [5 August 2025 11:06 UTC](https://cms-helpdesk.ncas.ac.uk/t/remote-setup-failing/1764/8 "2025-08-05T11:06:19Z")

</div>

I’ve tried to get u-dr157/run7 started again by getting it to poll. This is some of the workflow output:

> 2025-08-05T11:02:27Z INFO - Command “poll\_tasks” received. ID=6a10c32c-7f38-40c3-8cf3-c53b2e8b6233  
> poll\_tasks(tasks=[‘19801001T0000Z/atmos\_main’])  
> 2025-08-05T11:02:27Z INFO - Command “poll\_tasks” actioned. ID=6a10c32c-7f38-40c3-8cf3-c53b2e8b6233  
> 2025-08-05T11:02:28Z WARNING - platform: archer2-nvme - Could not connect to ln01.  
> \* ln01 has been added to the list of unreachable hosts  
> \* jobs-poll will retry if another host is available.  
> 2025-08-05T11:02:28Z WARNING - platform: archer2-nvme - Could not connect to ln02.  
> \* ln02 has been added to the list of unreachable hosts  
> \* jobs-poll will retry if another host is available.  
> 2025-08-05T11:02:29Z WARNING - platform: archer2-nvme - Could not connect to ln03.  
> \* ln03 has been added to the list of unreachable hosts  
> \* jobs-poll will retry if another host is available.

How do I fix?

For u-dr783 I just removed it and started again. That seems to be working… (Well jobs are started)

Simon

---

<div class="post-metadata">

**Author:** ![RosalynHatcher](https://dub1.discourse-cdn.com/flex013/user_avatar/cms-helpdesk.ncas.ac.uk/rosalynhatcher/32/11_2.png) [@RosalynHatcher](https://cms-helpdesk.ncas.ac.uk/u/RosalynHatcher)\
**Post date:** [5 August 2025 11:32 UTC](https://cms-helpdesk.ncas.ac.uk/t/remote-setup-failing/1764/9 "2025-08-05T11:32:43Z")

</div>

Hi Simon,

I’d probably try stopping and restarting the suite

Cheers,  
Ros.

---

<div class="post-metadata">

**Author:** ![SimonTett](https://avatars.discourse-cdn.com/v4/letter/s/3da27b/32.png) [@SimonTett](https://cms-helpdesk.ncas.ac.uk/u/SimonTett)\
**Post date:** [5 August 2025 12:13 UTC](https://cms-helpdesk.ncas.ac.uk/t/remote-setup-failing/1764/10 "2025-08-05T12:13:41Z")

</div>

Do I do that by:

cylc stop --now --now

and then cylc play ?

Simon

---

<div class="post-metadata">

**Author:** ![grenville](https://dub1.discourse-cdn.com/flex013/user_avatar/cms-helpdesk.ncas.ac.uk/grenville/32/29_2.png) [@grenville](https://cms-helpdesk.ncas.ac.uk/u/grenville)\
**Post date:** [5 August 2025 12:30 UTC](https://cms-helpdesk.ncas.ac.uk/t/remote-setup-failing/1764/11 "2025-08-05T12:30:12Z")

</div>

Simon

Yes, try that.

Grenville

---

<div class="post-metadata">

**Author:** ![SimonTett](https://avatars.discourse-cdn.com/v4/letter/s/3da27b/32.png) [@SimonTett](https://cms-helpdesk.ncas.ac.uk/u/SimonTett)\
**Post date:** [5 August 2025 14:49 UTC](https://cms-helpdesk.ncas.ac.uk/t/remote-setup-failing/1764/12 "2025-08-05T14:49:44Z")

</div>

It looks like 19801001T0000Z/pptransfer is hanging. Should I kill the task and trigger it?

Simon

---

<div class="post-metadata">

**Author:** ![grenville](https://dub1.discourse-cdn.com/flex013/user_avatar/cms-helpdesk.ncas.ac.uk/grenville/32/29_2.png) [@grenville](https://cms-helpdesk.ncas.ac.uk/u/grenville)\
**Post date:** [6 August 2025 07:29 UTC](https://cms-helpdesk.ncas.ac.uk/t/remote-setup-failing/1764/13 "2025-08-06T07:29:51Z")

</div>

The log says:

`[OK] Transfer command succeeded: globus transfer --format unix --jmespath ‘task_id’ --recursive --fail-on-quota-errors --sync-level checksum --label u-dr157/19801001T0000Z --verify-checksum --notify off 3e90d018-0d05-461a-bbaf-aab605283d21:/work/n02/n02/tetts/cylc-run/u-dr157/run7/share/cycle/19801001T0000Z a2f53b7f-1b4e-4dce-9b7c-349ae760fee0:/gws/nopw/j04/terrafirma/tetts/um_archive/u-dr157/19801001T0000Z`  
`[OK] Transfer: Transfer OK. (ReturnCode=0)`  
`2025-08-06T05:45:49Z INFO - succeeded`

I’d check on JASMIN that the data is all there, then if it is, change the task status to succeeded.

Grenville

---

<div class="post-metadata">

**Author:** ![system](https://europe1.discourse-cdn.com/flex013/uploads/cms_support/original/1X/1fd2411499ffcbc299fe756cd5cdf26e44956558.png) [@system](https://cms-helpdesk.ncas.ac.uk/u/system)\
**Post date:** [7 August 2025 07:30 UTC](https://cms-helpdesk.ncas.ac.uk/t/remote-setup-failing/1764/14 "2025-08-07T07:30:51Z")

</div>

This topic was automatically closed 24 hours after the last reply. New replies are no longer allowed.
