Skip to content

Fix/ocisdev 1433 - #748

Draft
2403905 wants to merge 5 commits into
stable-8.0from
fix/OCISDEV-1433
Draft

2403905 wants to merge 5 commits into
stable-8.0from
fix/OCISDEV-1433

Conversation

@2403905

@2403905 2403905 commented Sep 25, 2026

Copy link
Copy Markdown

No description provided.

2403905 and others added 5 commits September 24, 2026 20:36
…d in ListPublicShares (#735)

* fix(OCISDEV-877): skip public shares with nil resource_id in ListPublicShares

(cherry picked from commit 1ad0465)
Signed-off-by: Roman Perekhod <2403905@gmail.com>

* docs: [OCISDEV-877][stable-8.0] add changelog for nil resource_id fix

Signed-off-by: Roman Perekhod <2403905@gmail.com>

---------

Signed-off-by: Roman Perekhod <2403905@gmail.com>
Co-authored-by: Firas Frikha <firas.frikha@kiteworks.com>
The cs3 metadata storage authenticates as a system user before every
operation, and that Authenticate call carried no deadline. A gateway that
accepted the connection but never answered it - because it was itself
waiting on a stalled storage provider - parked the calling goroutine for
the lifetime of the process, with no log output at all.

The call now runs on its own bounded context and logs a warning when the
deadline is exhausted. The context returned to the caller deliberately
keeps no deadline: callers run their own RPC on it after getAuthContext
returns.

Adds the first tests for this package, including one that reproduces the
unbounded wait against an in-process gateway that never answers.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The grpc client connections of the pool were created without keepalive
parameters, so a peer that stopped answering on an established connection - a
black-holed node, a wedged process - was indistinguishable from a peer that was
merely slow, and a request without a deadline waited for the lifetime of the
process.

The clients now ping the peer while a request is in flight and fail the requests
on a connection that does not answer, tunable with GRPC_CLIENT_KEEPALIVE_TIME
and GRPC_CLIENT_KEEPALIVE_TIMEOUT. The servers got the matching enforcement
policy so they accept those pings.

GRPC_MAX_CONNECTION_AGE is removed along with it. It closed healthy connections
on a timer, never ended a request that was already in flight because the grace
period was left at infinity, and fell back to doing nothing at all whenever its
value had no unit suffix.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@kw-security

kw-security commented Sep 25, 2026 •

Copy link
Copy Markdown

✅ Snyk checks have passed. No issues have been found so far.

Status Scan Engine Critical High Medium Low Total (0)
✅ Open Source Security 0 0 0 0 0 issues
✅ Licenses 0 0 0 0 0 issues
✅ Code Security 0 0 0 0 0 issues

💻 Catch issues earlier using the plugins for VS Code, JetBrains IDEs, Visual Studio, and Eclipse.

@kobergj kobergj left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reviewed the public-share manager changes (d4f339c + 95d977d). The direction is right - taking the read-modify-write out of GetPublicShare / GetPublicShareByToken / ListPublicShares, dropping the manager lock before the gateway Stat fan-out, and deleting the signal.Notify(SIGHUP, SIGINT, SIGQUIT) janitor hack (which was suppressing the default signal disposition process-wide) are all clear improvements. Wiring Close through rgrpc.cleanupServices and passing the request context into init() are good too.

One blocking issue in cs3.Write (cache/mtime inconsistency that defeats IfUnmodifiedSince), plus a few smaller things inline.

One process note: no changelog entry for these two commits, although the other commits in this PR each add one - and this change is user-visible (expired shares are no longer purged on access, janitor cadence changed, new shutdown semantics).

// independent copy (see persistence.Copy), it has to be done explicitly
// here, or the cache would only pick up our own write once some later
// external write advances the remote mtime past our stale one.
if info, statErr := p.s.Stat(ctx, "publicshares.json"); statErr == nil {

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Blocking. Pairing our content with an mtime from a separate Stat after the upload makes the cache self-inconsistent, and that defeats the IfUnmodifiedSince guard on line 139.

Before this change, p.db.mtime was only ever assigned in Read, right next to the content downloaded at that mtime - they were consistent by construction. With two instances writing the same publicshares.json:

  1. A: Read -> cache mtime M0
  2. A: Upload(C_A, IfUnmodifiedSince: M0) -> OK, remote becomes M1
  3. B, in the window between A's PUT and A's Stat: Read -> M1, Upload(C_B, IfUnmodifiedSince: M1) -> OK, remote becomes M2/C_B
  4. A: Stat -> M2. A now caches mtime M2 with content C_A
  5. A's next Read: M2.After(M2) == false -> no refetch -> A keeps serving C_A, so B's share is invisible on A indefinitely
  6. A's next Write: IfUnmodifiedSince: M2, remote mtime is M2, After is false -> precondition passes -> A uploads C_A + its mutation, silently dropping B's share

UploadResponse only carries Etag/FileID, so there is no mtime to reuse. Two ways out:

  • Invalidate instead of syncing: on success set p.db.mtime = time.Time{} and clear the cached content, so the next Read unconditionally refetches. One extra download per write, obviously correct.
  • Or set MTime on the UploadRequest (the metadata storage supports it via X-OC-Mtime) and cache that value. Content and mtime stay consistent, and it removes this Stat - a full round trip currently held under p.mu.

Either way, statErr should not be swallowed silently.

for _, v := range db {
var changed bool
for id, v := range db {
d := v.(map[string]interface{})["share"]

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

These two assertions still panic: v.(map[string]interface{}) and d.(string). This runs in a bare goroutine with no recover(), so one malformed entry in publicshares.json takes the whole service down - and repeats every janitor tick.

The continue added just below shows the intent was to tolerate bad entries, so it seems worth extending it one line up, especially since OCISDEV-877 in this same PR exists because partially-broken entries do occur:

d, ok := v.(map[string]interface{})["share"].(string)
if !ok {
    continue
}

var ps link.PublicShare
if err := utils.UnmarshalJSONToProtoV1([]byte(d), &ps); err != nil {
    continue
}

return err
}

m.mutex.Lock()

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This holds the exclusive manager lock across a persistence Read and Write - two metadata-storage round trips, bounded at 60s - blocking every public-share read in the process for the duration.

Much better than the previous N sequential read+write cycles under the lock, so not blocking. But for a change whose point is read-path contention, the janitor is now the remaining stop-the-world point. Worth considering: collect the expired IDs under RLock, then take the write lock only for the read-modify-write, and let IfUnmodifiedSince + a retry handle the race.

// that a caller which keeps reading the result after releasing its lock
// cannot race a writer that later mutates an existing share's fields in
// place (see manager.UpdatePublicShare).
func Copy(db PublicShares) PublicShares {

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This now runs on every read path, so each ListPublicShares / GetPublicShareByToken allocates a fresh copy of the entire share database (N+1 maps). That is a real cost on instances with a large publicshares.json, which are exactly the ones hurting from read-path contention today.

Note that the cs3 cache is already effectively immutable once published: both the refill in Read and Write replace p.db.publicShares with a fresh map rather than mutating it in place. The copy is only needed because the manager's write paths mutate what Read handed them - UpdatePublicShare doing data["share"] = ..., and delete(db, id) in revokePublicShare / cleanupExpiredShares.

So Copy could move into those four write paths and Read could document its result as read-only. Same race fix, allocation-free reads.

}
if c.JanitorRunInterval == 0 {
c.JanitorRunInterval = 60
c.JanitorRunInterval = 600

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

60 -> 600 is a 10x change to the default cleanup cadence, and it is not mentioned in the commit message or a changelog. Intentional? It is defensible now that the read paths no longer purge, but on a stable branch it should be deliberate and documented.


const writers = 8
const readers = 8
const iterations = 500

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Nit: 8 writers x 500 iterations is 4000 real disk writes plus 4000 reads - ~15s locally without -race, and noticeably worse in CI with it. ~50 iterations pins the same regression.

// Overlapping reads should finish close to a single delay - allow
// generous slack for scheduling noise without letting a real
// regression pass.
Expect(time.Since(start)).To(BeNumerically("<", delay*3))

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Nit: this asserts on a wall-clock ratio (< 450ms vs 1.2s serialized), which will be flaky on a loaded shared runner. The property being tested is the right one - a counter in slowReadPersistence.Read tracking max observed concurrency (Expect(maxConcurrent).To(BeNumerically(">", 1))) would test it deterministically.


// Give the Read goroutine time to be inside its (slow) Stat call,
// holding mu, before Init races it.
time.Sleep(delay / 5)

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Nit: time.Sleep(delay / 5) to get the other goroutine inside its Stat, then asserting < delay/2, is the same wall-clock flakiness as the ginkgo test in json_test.go. Signalling from inside the fake Stat (a channel the storage closes on entry) removes the guess.

return m, nil
}

var _ publicshare.ClosableManager = (*manager)(nil)

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Nit: this interface assertion reads better next to the manager type declaration than between New and commonConfig.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants