It seems the majority of tests status is "queued", but their environment status table still has a result. When attempting to manually re-test I get a fatal error:

Fatal error: require_once() [function.require]: Failed opening required '/.server.inc' (include_path='.:/usr/share/php5:/usr/share/php') in /var/www/qa.drupal.org/htdocs/sites/all/modules/project_issue_file_review/server/pifr_server.review.inc  on line 151

I'm going to assume the fact that they get marked for re-test so many times and never tested is due to that.

Comments

boombatower’s picture

It seems the tests experiencing this never had any result returned...yet they show mysql as failed...and fail to queue for retesting since they are queued.

boombatower’s picture

StatusFileSize
new63.39 KB

To demonstrate how weird this is behaving.

client test

boombatower’s picture

I cannot re-create these problems on qa-scratch.

webchick’s picture

Priority: Normal » Critical

This has completely blocked core development for the past three days. I think that qualifies as critical.

What can we do to get you some help to debug this?

boombatower’s picture

#705290: Database dump

I really have no idea what's going on. The symptoms aren't even consistent, much less the fact that I can do the same thing on qa-scratch and qa and it will fail on qa and work on qa-scratch.

dave reid’s picture

Yeah this is really wierd. Seems like we're stuck somewhere in limbo with confirming client/test slave tests. Something's obviously up since this has been a couple days now.

pwolanin’s picture

Would it be possible to install fresh and queue tests starting from some known date?

boombatower’s picture

That or just fry all data over last week to two weeks. I am really not sure what else to do.

rfay’s picture

IMO it would be better to just get things going again one way or another, if that's possible, even it it means using a fresh clean box or something.

boombatower’s picture

StatusFileSize
new12.31 KB

With my read access to qa.drupal.org database I have discovered some things.

1. pifr_test_environment looks good (so relevant environment determination code works)
2. results get environment_id of 0
3. why ignored re-test requests do not send back a positive response to project client (unrelated)

My assumption is something has to be messed up with result saving to generate a 0 entry, so lets look at where results are saved.

boombatower@boomba:~/software/project_issue_file_review> grep -nR --exclude-dir=CVS "pifr_server_result_save" ../server/pifr_server.result.inc:148:function pifr_server_result_save(array $result) {
./server/pifr_server.test.inc:352:  $result = pifr_server_result_save($result);
./pifr.install:316:      pifr_server_result_save($test_result);

pifr_server.test.inc:352

$result['environment_id'] = pifr_server_environment_status_get_client($test, $client);

/**
 * Get the environment ID that the client is currently testing.
 *
 * @param array $test Test information.
 * @param array $client Client information.
 * @return interger Environment ID, or FALSE.
 */
function pifr_server_environment_status_get_client(array $test, array $client) {
  $result = db_query('SELECT environment_id
                      FROM {pifr_environment_status}
                      WHERE test_id = %d
                      AND client_id = %d', $test['test_id'], $client['client_id']);
  return db_result($result);
}

So that leaves us with a) the pifr_environment_status tables gets messed up, b) is magically empty and FALSE is returned and cast to int as 0, or c) magic!

All values for a fresh test that was only queued, never re-queued, never tested look good.

Interestingly all the results that have environment_id of 0 are of code 3 or 5 (only 2 5's). 3 being a failure to checkout code. This would seem like someone had a wonky client running (which there were several news one recently), and it hosed everything.

boombatower’s picture

Looking at the log for a test that got messed up you see the following.

636,776 	Re-test request ignored because test is already queued.	02/01/2010 - 15:00:47
636,726 	Re-test request ignored because test is already queued. 	02/01/2010 - 14:50:57
636,684 	Re-test request ignored because test is already queued. 	02/01/2010 - 14:40:26
636,642 	Re-test request ignored because test is already queued. 	02/01/2010 - 14:30:44
636,594 	Re-test request ignored because test is already queued. 	02/01/2010 - 14:20:32
636,500 	Re-test request ignored because test is already queued. 	02/01/2010 - 14:10:28
636,456 	Re-test request ignored because test is already queued. 	02/01/2010 - 14:00:36
636,086 	Result received from test client #33. 	02/01/2010 - 12:58:35
636,084 	Requested by test client #33. 	02/01/2010 - 12:58:33
636,082 	Test reset by client request. 	02/01/2010 - 12:58:33
636,080 	Requested by test client #33. 	02/01/2010 - 12:58:33 

Test client #33 is the only one with two environments: mysql and coder. Scratch has no clients that support both (registered that way anyway), but does have the coder environment enabled. Thus the only decent difference between the setups.

boombatower’s picture

The place where all the code starts is:

/**
 * Get the next eligible branch or file test.
 *
 * @param array $client Client information.
 * @return array Test information and environment or FALSE.
 */
function pifr_server_test_get_next_eligible(array $client) {
  $result = db_query('SELECT z.*
                      FROM
                      (
                        SELECT t.test_id, e.environment_id
                        FROM {pifr_test} t
                        JOIN {pifr_test_environment} te
                          ON t.test_id = te.test_id
                        JOIN {pifr_environment} e
                          ON (e.environment_id = te.environment_id AND e.environment_id IN (' . db_placeholders($client['environment'], 'int') . '))
                        WHERE t.type != %d
                        AND t.status = %d
                        ORDER BY t.type, t.test_count, t.last_tested, t.test_id
                      ) z
                      LEFT JOIN {pifr_environment_status} s
                        ON (z.test_id = s.test_id AND z.environment_id = s.environment_id)
                      LEFT JOIN {pifr_result} r
                        ON (z.test_id = r.test_id AND z.environment_id = r.environment_id)
                      WHERE s.client_id IS NULL
                      AND r.result_id IS NULL
                      LIMIT 1',
    array_merge($client['environment'], array(PIFR_SERVER_TEST_TYPE_CLIENT, PIFR_SERVER_TEST_STATUS_QUEUED)));

  // If result found then load test and environment.
  if ($info = db_fetch_array($result)) {
    $test = pifr_server_test_get($info['test_id']);
    $environment = pifr_server_environment_get($info['environment_id']);
    return array($test, $environment);
  }
  return FALSE;
}

thus I will setup a client on qa-scratch just like #33 and then see if I can replicate the 0 somewhere in the code.

boombatower’s picture

Setup coder review on #32 which I have access to and got everything running on qa-scratch...running client confirmation at the moment.

boombatower’s picture

Well HEAD is broken, so I put DamZ's client into debug mode so it would only run one test and pass...worked. I then ran a number of patches through the system..not problems.

I am really starting to think a data flush is in order or at least more recent data. Something is very weird....and I just can't seem to track it down to recreate it.

I'll try and be around tomorrow so we can make a decision, its 4am here.

berdir’s picture

FYI, all tests except page cache and update core/contrib with the patch at #705854: clickLink() does not work for url_target containing a fragment and breaks testing worked for me and the failing tests could be due to my setup, I've seen atleast the page cache fails before.

Without the linked patch, there are over 105 fails, and 27 exceptions.

boombatower’s picture

Decided to just fry bad data, I ask killies to run the following:

UPDATE pifr_test SET status = 2 WHERE test_id IN (SELECT test_id FROM pifr_result WHERE environment_id = 0); # should be 61
DELETE FROM pifr_result WHERE environment_id = 0; # should be 61
DELETE FROM pifr_client_test; # should be 13
DELETE FROM pifr_result WHERE test_id IN (27318,26762,27100,27274,27300,22912,26922,27228,27048,27232,23146,16784,26854,27216,27218,27224,27220,27246,27314,27254,27242,27214,27256,27312,27192,27306,27304,27308,27310,27316,27126,27330,27328,26130,26510,27298,22906,27244,27238,27268,27234,27248,27258,27260,27250,27194,27196,27208,27204,27202,27198,27206,27210,27212,27190,27188,27184,27178,27174,27176,27180);

Query OK, 61 rows affected (0.03 sec)
Query OK, 61 rows affected (0.03 sec)
Query OK, 13 rows affected (0.00 sec)
Query OK, 60 rows affected (0.00 sec)

Better query if we have to do this again (hopefully after figure it out):

UPDATE pifr_test SET status = 2 WHERE test_id IN (SELECT test_id FROM pifr_result WHERE environment_id = 0);

SELECT test_id FROM pifr_result WHERE environment_id = 0;
DELETE FROM pifr_result WHERE test_id IN ( [LIST FROM PREVIOUS QUERY] );

DELETE FROM pifr_client_test;
boombatower’s picture

Status: Active » Needs review
StatusFileSize
new1.74 KB

DamZ found the piece of flawed logic (oversight from addition of environments). I have updated logic in patch.

boombatower’s picture

Title: Discover weird mis-queue issue » Correct pifr.result() validation
Status: Needs review » Fixed

Committed.

Status: Fixed » Closed (fixed)

Automatically closed -- issue fixed for 2 weeks with no activity.