Hi Chris and Noah, last year we exempted type strain records from outlier filtering, which has cleaned up many blasts. Now I'm collecting the instances where legitimate type strain records are being excluded, and trying to learn from these instances how the logic could be further modified toward the goal of including true type records and filtering incorrectly classified type records.
Here are 3 examples:
I think excluding genomes that fail taxonomy check is low risk and that if possible we should implement that. This implementation would also clean up some other blasts, for example excluding LSCY01000004 and LSCY01000011 here.
For all 3 of these instances, all type strain records in the db end up in the outlier pile. Can we lean on that as a criterion to exempt a taxon from type strain outlier filtering?
Hi Chris and Noah, last year we exempted type strain records from outlier filtering, which has cleaned up many blasts. Now I'm collecting the instances where legitimate type strain records are being excluded, and trying to learn from these instances how the logic could be further modified toward the goal of including true type records and filtering incorrectly classified type records.
Here are 3 examples:
Faecalibacterium longum https://gitlab.labmed.uw.edu/molmicro/uwnt-management/-/issues/359
Segatella baroniae https://gitlab.labmed.uw.edu/molmicro/uwnt-management/-/issues/361
Corynebacterium pseudogenitalium https://gitlab.labmed.uw.edu/molmicro/uwnt-management/-/issues/371
I think excluding genomes that fail taxonomy check is low risk and that if possible we should implement that. This implementation would also clean up some other blasts, for example excluding LSCY01000004 and LSCY01000011 here.
For all 3 of these instances, all type strain records in the db end up in the outlier pile. Can we lean on that as a criterion to exempt a taxon from type strain outlier filtering?