*ttl drivers: a dead TTL fiber cannot be restarted and wedges the queue in ENDING on the next RO switch
Avaliação
- Dificuldade
- 4/5
- Tempo estimado
- 3-5 dias
- Facilidade para iniciantes
- 48/100
Direção de pesquisa
Start with queue/abstract/driver/fifottl.lua lines 164-167 and 415-427, then compare the buffered-channel behavior in utubettl. Run the reproduction in the issue and inspect queue_state handling during RO/RW switches. Done means a failed TTL iteration does not permanently disable processing or leave the queue stuck in ENDING, and start/stop remain safe.
Escrita pelo modelo de indexação a partir do texto da issue.
Descrição
Environment: queue master (07fd732), Tarantool 3.8.0 (also the same code in 1.5.0). memtx.
Related: #238 asks for an API to restart the TTL fiber. This issue is about what happens today when the fiber dies: it cannot be restarted, and for fifottl/limfifottl the queue state machine gets stuck in ENDING on the next RO switch, so the queue never comes back without an instance restart.
Summary
Any error inside a TTL iteration (a user on_task_change callback raising, an MVCC conflict, a vinyl read error, the nil dereference from the TTL branch, ...) makes the fiber return 1 (fifottl.lua#L164-L167) while self.fiber keeps pointing at the dead fiber. Consequences:
method.start()is a no-op becauseself.fiberis not nil (fifottl.lua#L415-L420). TTL/TTR/delay processing for the tube is dead.- On the next RO switch the state machine calls
stop(), which doesself.sync_chan:put(true)(fifottl.lua#L427). The channel is unbuffered and the only reader is dead (same root cause as #262), so thequeue_statefiber blocks forever. The queue stays inENDING, andput/takekeep failing withqueue is in ENDING stateeven after the instance is RW again.
For utubettl (buffered fiber.channel(1)) point 2 does not hang, but point 1 still holds until an RO->RW cycle.
Repro
local fiber = require('fiber')
box.cfg{}
local queue = require('queue')
local fired = false
local tube = queue.create_tube('t', 'fifottl', {
on_task_change = function(task, stat)
if stat == 'ttl' and not fired then
fired = true
error('user callback failed once')
end
end,
})
tube:put('a', {ttl = 0.1})
fiber.sleep(0.3)
print('1. fiber after one callback error:', tube.raw.fiber:status())
local b = tube:put('b', {ttl = 0.1})
fiber.sleep(0.3)
print('2. expired task still in space:', tube.raw.space:get(b[1]) ~= nil)
tube.raw:start()
print('3. fiber after start():', tube.raw.fiber:status())
box.cfg{read_only = true}
fiber.sleep(1)
print('4. queue.state() after RO switch:', queue.state())
box.cfg{read_only = false}
fiber.sleep(1)
print('5. queue.state() after RW switch:', queue.state())
print('6. put() after RW switch:', pcall(tube.put, tube, 'c', {ttl = 0.1}))
Output on master:
1. fiber after one callback error: dead
2. expired task still in space: true
3. fiber after start(): dead
4. queue.state() after RO switch: ENDING
5. queue.state() after RW switch: ENDING
6. put() after RW switch: true nil -- put() returns nil, log: "put: queue is in ENDING state"
Expected
A failed iteration should not permanently kill TTL processing: either the fiber survives the error (log it and continue / restart with backoff), or at least self.fiber is cleared on exit so start() can create a new one, and stop() must never block on a channel nobody reads.
Suggested fix
- Clear
self.fiberin a guaranteed cleanup path when the fiber exits (checking it still points to the current fiber). - Make
stop()tolerant of a dead fiber (checkself.fiber:status()or use a buffered channel /fiber:cancel()). - Optionally restart the fiber after an error with a backoff instead of
return 1; user callback errors could be pcall-ed inabstract.luaso a user bug cannot take the driver down.
- Linguagem predominante
- Lua
- Estrelas
- 243
- Forks
- 56
- Merge médio
- 6d 8h
- PRs com merge (30d)
- 1
Preparar o ambiente
Este projeto não oferece contêiner de desenvolvimento, Dockerfile nem guia de contribuição, então a configuração fica por sua conta: comece pelo README e veja nosso guia da primeira contribuição para os passos gerais.
Primeiros passos
- Leia a issue inteira e depois o guia de contribuição do projeto.
- Comente na issue dizendo que vai assumir — evita que duas pessoas façam o mesmo trabalho.
- Faça um fork do repositório e trabalhe em uma branch.
- Abra um pull request que referencie o número da issue.
Mais de tarantool/queue
-
documentation good first issue
Dificuldade 1/5 Menos de uma hora Facilidade para iniciantes 85/100
-
*ttl drivers: TTL branch dereferences delete() result without a nil check, killing the fiberTalvez já em andamento @maksimuimin assumiu há 14 dias. Aberta
Dificuldade 3/5 1-2 dias Facilidade para iniciantes 76/100
-
fifottl/limfifottl: stop() blocks forever while RW, drop() leaks the TTL fiberTalvez já em andamento @maksimuimin assumiu há 14 dias. Aberta
Dificuldade 4/5 3-5 dias Facilidade para iniciantes 55/100
-
1sp bug teamE
Dificuldade 4/5 3-5 dias Facilidade para iniciantes 25/100
-
bug
Dificuldade 4/5 3-5 dias Facilidade para iniciantes 35/100
Todas as issues de tarantool/queue
Issues semelhantes
-
Dificuldade 2/5 1-3 horas Facilidade para iniciantes 85/100
cataclysmbn/Cataclysm-BN#10516 ·
Mantenedores costumam responder em até 1 dia
-
Dificuldade 2/5 1-3 horas Facilidade para iniciantes 68/100
silverbulletmd/silverbullet#2187 ·
Mantenedores costumam responder em até 2 dias
-
clangd_extensions.nvim and alabaster.nvim: "p00f" username now belongs to a different account (repo-jacking risk)Talvez já em andamento Um pull request vinculado a esta issue está aberto ou já foi mesclado. Aberta
Dificuldade 2/5 1-3 horas Facilidade para iniciantes 85/100
AstroNvim/astrocommunity#1802 ·
-
Bug era/hc Miscellaneous Task
Dificuldade 2/5 1-3 horas Facilidade para iniciantes 72/100
Mantenedores costumam responder em até 1 dia
-
Dificuldade 2/5 1-3 horas Facilidade para iniciantes 64/100