The p-value is NOT the probability that B is better: it is the probability of seeing a difference like this if there were none. And it only holds if you fixed the sample size beforehand and didn't stop on seeing a good result (“Peeking early” tab).
Enter the visitors and conversions of each variant (A is the control). First the traffic split is checked (SRM): if it doesn't match the requested one, the result isn't reliable and you need to find the cause. Then you get the difference with its confidence interval, which tells you how much, not just whether. With little data or a rate at 0% or 100% Fisher's exact test is used, because the normal approximation declares winners that aren't.
The p-value is the probability of seeing a difference like this or larger if there were really none. It is not the probability that B is better: the Bayesian part gives that, P(B > A). With more than two variants it is corrected for multiple comparisons (Holm or Benjamini-Hochberg): comparing several variants at once gives more chances to win by luck.
The p-value is NOT the probability that B is better: it is the probability of seeing a difference like this if there were none. And it only holds if you fixed the sample size beforehand and didn't stop on seeing a good result (“Peeking early” tab).