Skip to content
Merged
Show file tree
Hide file tree
Changes from 3 commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 1 addition & 1 deletion checkstyle/suppressions.xml
Original file line number Diff line number Diff line change
Expand Up @@ -161,7 +161,7 @@
files="StreamThread.java"/>

<suppress checks="ClassDataAbstractionCoupling"
files="(KStreamImpl|KTableImpl).java"/>
files="(KafkaStreams|KStreamImpl|KTableImpl).java"/>
Comment thread
ableegoldman marked this conversation as resolved.
Outdated

<suppress checks="CyclomaticComplexity"
files="(StreamsPartitionAssignor|StreamThread|TaskManager|PartitionGroup).java"/>
Expand Down
44 changes: 33 additions & 11 deletions streams/src/main/java/org/apache/kafka/streams/KafkaStreams.java
Original file line number Diff line number Diff line change
Expand Up @@ -92,6 +92,7 @@
import java.util.Set;
import java.util.TreeMap;
import java.util.UUID;
import java.util.concurrent.atomic.AtomicInteger;
import java.util.concurrent.ExecutionException;
import java.util.concurrent.Executors;
import java.util.concurrent.ScheduledExecutorService;
Expand Down Expand Up @@ -463,9 +464,8 @@ private void replaceStreamThread(final Throwable throwable) {
closeToError();
}
final StreamThread deadThread = (StreamThread) Thread.currentThread();
threads.remove(deadThread);

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I remember that we had the replace use the same ID for a reason. (maybe it had to do with rebalancing?). I don't think there should be a problem to try to get the same ID by waiting a bit in the replace thread

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We cannot wait here until the dead thread is shutdown because the shutdown happens after replaceStreamThread() throws the exception. So we would wait forever.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

deadThread.shutdown(); I was referring to this below. But if we don't need to keep the same for any reason I am fine either way

@cadonna cadonna Mar 1, 2021

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Sorry for not being clear enough. I was also referring to deadThread.shutdown(). Method deadThread.shutdown() only requests a shutdown. The actual shutdown is performed in completeShutdown() which is called after replaceStreamThread() throws one of the exceptions below. Since completeShutdown() is called by the same thread that calls deadThread.shutdown() in this method, i.e., the dead thread, we would wait forever if we waited after deadThread.shutdown() until the dead thread is shut down. Or did I misunderstand your statement "by waiting a bit in the replace thread"?

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Ah that makes sense. Another approach is using a temporary name. And then waiting the new thread until the old thread dies and takes the name. This is a bit complicated and I think it should only be done if it is necessary for the new thread to have the same name. And probably not in this PR but it could be an improvement done later

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yeah I think swapping the names would make the code unnecessarily complicated, and it would definitely make reading the logs more difficult.
Just to note: in the current rebalance protocol, the thread name should not impact the task assignment since within a client tasks are always just assigned to their previous owner (we maximize stickiness & balance)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

since the name doesn't matter I wonder why we spent so much effort making sure it had the same name?

addStreamThread();
deadThread.shutdown();
addStreamThread();
if (throwable instanceof RuntimeException) {
throw (RuntimeException) throwable;
} else if (throwable instanceof Error) {
Expand Down Expand Up @@ -1047,9 +1047,15 @@ private Optional<String> removeStreamThread(final long timeoutMs) throws Timeout
if (!streamThread.waitOnThreadState(StreamThread.State.DEAD, timeoutMs - begin)) {
log.warn("Thread " + streamThread.getName() + " did not shutdown in the allotted time");
timeout = true;
// Don't remove from threads until shutdown is complete. We will trim it from the
// list once it reaches DEAD, and if for some reason it's hanging indefinitely in the

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Where do we trim this list? I don't thing we do. In the begging of addStreamThread() can we purge the dead threads? That is the only place it should matter

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We're trimming it in getNextThreadIndex. But if we're going to rely on threads.size() elsewhere, which it seems we do, then yeah we should trim it more aggressively

// shutdown then we should just consider this thread.id to be burned
} else {
threads.remove(streamThread);

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

If we purge the dead threads before we add new ones and if we remove the assumption that there are no dead threads in the thread list we can just not remove the threads in remove thread. This will make it there should be no concern about the cache size changing when a thread is removing itself. And make the risk we took about memory overflows unnecessary.

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yeah, I was mainly trying to keep things simple. There's definitely a tradeoff in when we resize the cache: either we resize it right away and risk an OOM or we resize it whenever we find newly DEAD threads but potentially have to wait to "reclaim" the memory of a thread.
Both scenarios run into trouble when a thread is hanging in shutdown, but if that occurs something has already gone wrong so I don't think we need to guarantee Streams will continue running perfectly. But the downside to resizing the cache only once a thread reaches DEAD is that a user could call removeStreamThread() with a timeout of 0 and then never call add/remove thread again, and they'll never get back the memory of the removed thread since we only trim the threads inside these methods (or the exception handler). ie, it seems ok to lazily remove DEAD threads if we only use the threads list to find a unique threadId, but not to lazily resize the cache. WDYT?

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think we can leave it for now, if we should see problems this could be a fix, we don't run a single thread soak so we won't see this issue ourselves but there are many single thread applications that could start using this and we should see if they have problems

}
}
threads.remove(streamThread);
// Don't remove from threads until shutdown is complete since this will let another thread
// reuse its thread.id. We will trim any DEAD threads from the list later
final long cacheSizePerThread = getCacheSizePerThread(threads.size());

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We need to adapt the resizing of the cache per threaad to use only the number of non-DEAD stream threads instead of all stream threads in the list. There are other two locations where we use the size of the thread list to resize the cache per thread.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Perviously we had relied on the fact there were no dead threads in the list

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Good catch. I added a method getNumLiveStreamThreads to use instead of just threads.size() which will trim the list of any DEAD threads and return the actual number of living threads

resizeThreadCache(cacheSizePerThread);
if (groupInstanceID.isPresent() && callingThreadIsNotCurrentStreamThread) {
Expand Down Expand Up @@ -1094,16 +1100,32 @@ private Optional<String> removeStreamThread(final long timeoutMs) throws Timeout
}

private int getNextThreadIndex() {
final HashSet<String> names = new HashSet<>();
processStreamThread(thread -> names.add(thread.getName()));
final String baseName = clientId + "-StreamThread-";
for (int i = 1; i <= threads.size(); i++) {
final String name = baseName + i;
if (!names.contains(name)) {
return i;
final HashSet<String> allLiveThreadNames = new HashSet<>();
final AtomicInteger maxThreadId = new AtomicInteger(1);

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Do we really need an atomic integer here? maxThreadId is only used in the synchronized block.

@ableegoldman ableegoldman Mar 1, 2021

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It's because of the whole "variables used in a lambda must be final or effectively final" thing

synchronized (threads) {
processStreamThread(thread -> {
// trim any DEAD threads from the list so we can reuse the thread.id
// this is only safe to do once the thread has fully completed shutdown
if (thread.state() == StreamThread.State.DEAD) {
threads.remove(thread);
} else {
allLiveThreadNames.add(thread.getName());
final int threadId = thread.getName().charAt(thread.getName().length() - 1);
Comment thread
ableegoldman marked this conversation as resolved.
Outdated
if (threadId > maxThreadId.get()) {
maxThreadId.set(threadId);
}
}
});

final String baseName = clientId + "-StreamThread-";
for (int i = 1; i <= maxThreadId.get(); i++) {
final String name = baseName + i;
if (!allLiveThreadNames.contains(name)) {
return i;
}
}
return threads.size() + 1;
}
return threads.size() + 1;
}

private long getCacheSizePerThread(final int numStreamThreads) {
Expand Down