Hacktoberfest 2026: the issues maintainers tagged for October, open and beginner-friendly. Browse Hacktoberfest issues

diff_analysis with multiple groups

Open
#110 0 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

Assessment

Difficulty
4/5
Estimated time
3-5 days
Newbie friendliness
25/100
Issue type
Documentation
Clarity
Needs clarification
Activity status
Stale
Tech stack
r

Research direction

Start by reading the diff_analysis entry point and its handling of the lda, kruskal.test, and wilcox.test options. Determine why four-group and A/B-only comparisons produce different discriminative features, and clarify whether separate pairwise comparisons are expected. Done means documenting the behavior and recommended comparison approach.

Written by the indexing model from the issue text.

Description

Hi,

I was hoping to get a bit more understand of the diff_analysis, if possible. I am struggling to understand why I get different taxa as significant from the diff_analysis function if I compare my 4 groups vs if I subset the data and compare 2 of the groups.

I have a class group with 4 groups - A, B, C, D. When I use diff_analysis I get only 3 discriminative features after lda.

If I subset the class group to only have options A, B - when I use diff_analysis I get 78 features discriminative features after lda.

The code used was:

set.seed(50)
deres <- diff_analysis(obj = ps_sub_ran_for_tree, classgroup = "cluster",
mlfun = "lda",
filtermod = "pvalue",
firstcomfun = "kruskal.test",
firstalpha = 0.05,
strictmod = TRUE,
secondcomfun = "wilcox.test",
subclmin = 10,
subclwilc = TRUE,
secondalpha = 0.05,
lda=2)

The results when comparing just A and B match a lot of the ones found using random forest, while comparing A,B,C,D does not. Should I be doing each pairwise comparison separately?

Thank you for you help!

Dominant language
R
Stars
195
Forks
36
PR merge metrics
No merged PRs in 30d

Getting set up

This project ships no dev container, Dockerfile or contributing guide, so setting up is up to you: start from its README, and see our first-contribution guide for the general steps.

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

More from YuLab-SMU/MicrobiotaProcess

All issues in YuLab-SMU/MicrobiotaProcess

Similar issues

More R issues

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.