BackBurner Probleme...

Anyone have experienced problems between BackBurner and Vray?

When I render on the renderfarm. Suddenly the manager stop sending scenes to some servers. Without any warning or error message, the manager suddenly “forget” some server.

After few minutes or few hours only 2 or 3 nodes still rendering. The others nodes apears Yellow (idle) and no error messages apear. I have verify and it’s not because Backburner have already assign the working node to the remaining jobs or frame to renders…

I’m not sur but this bug seem happening only on heavy or long render. I never saw this problem before.

Thank you for your help.

Did they throw errors before ? (check the error tab of the corresponding job). Backburner automatically pulls servers from job after a certain amount of errors.

Regards,
Thorsten

Nope… No error at all…

Backup Burner at end only keep one render Node active. And has I said, the remaining frames are not already assign…

It’s really a strange behaviour…

I appologize for my bad english…

That sounds like a problem we had awhile back. There were plenty of jobs in the queue, but not all the servers were kicking in. What version of BB are you using? I think that happened for us with the version that shipped with max9. There was a service pack for that one which fixed it.

I can’t remember completely, but some people fixed problems with it by reverting only the manager back to an older version.

Maybe I can try that… But I don’t have a lot of time to put on a complete BackBurner re-install of the renderfarm…

We have the last service pack…

I will try to re-install it first on de manager Node… I will se after…

Thank You!

I have the exact same problem and I’ve got the latest SP installed.

We recently set our manager up on a 64-bit machine (to get around the 600 minute task timeout issue), and the first job I rendered over the weekend had the same problem you describe. About 20% of the render farm just went idle and refused to be reassigned. I read on the Area forum that some people got around it by suspending the job and restarting, but I haven’t had a chance to confirm that yet.

I never had this particular problem with a 32-bit manager. Only other problems.

I’m running Max 9 32bit so it’s not just limited to the 64 bit version, and I can confirm that restarting the job usually fixes the problem at least on that job.

It’s not doing anything weird like allocating blocks of frames is it? It’s a default preference in the submit dialog of backburner under advanced where it’ll assign 10 frames to one machine, ten to another and so on rather than each frame going to the next free machine.

No it will stop sending frames to where only one or two machines are actively rendering when there are hundreds of frames left.

Anyone got any other tips for these backburner issues? We get Application Load Timeouts occurring regularly, but 95%-99% of the time these only occur on dual quadcore machines (running XP64). Its really frustrating as I set a whole lot of jobs rendering overnigh/overweekend and come back to find that our fastest machines have been idle for most of the time!

We are using Max2009 64bit SP2 on all machines with Vray 1.5 SP2 and I am rendering the frames from a saved lightcache solution (but the primary bounce is brute force).

BB is version 2008.1

When I look at the BB Server screen on the offending machines, I get the message:

Server has not received an update from the manager and has timed out; assuming the manager is down.
Manager is not responding.

I can’t be sure that this happens all the time, but on checking this morning, it vertainly seems to be the case.

Me again! The BB Manager program runs from our server which is on Windows Server 2003. It is not 64bit. Could this be an issue?

Our manager is run on Windows Server 2003 32-bit and no problems here. The only time we get the “manager down” message is when we cancel a job or remove a render node from a particular job, then it freaks out a little. This happens mostly on our dual quadcore machines (as you noted) however, it will pick back up after a few minutes.

We had the application load timeout when we first started building a farm because our machines were to slow to load the file before the manager gave up on them. However, that shouldn’t happen on your machines. How much ram do they have? For a dual quad, 4 gb is the minimum. Treat it like each core is its own system. With 4 gigs, that means you have 8 cores with 512K ram each. This has been fine for us, but we do have scenes on occasion when these systems render slower than they should because of this. In reality, 8gb of ram would be better and we will probably upgrade soon.

If I remember right, there was a setting somewhere that set how long the manager would wait for the server to load the scene and start rendering. I don’t know if it’s in the job settings when you submit the job or in the manager settings. You also might look in the backburner.xml in the network folder in the backburner install directory on the manager system. There might be a setting in there to try.

It could also have something to do with your network speed. Remember, that the manager has to distribute the file over the network to all the nodes and if it takes too long, the manager might give up on them.

That’s all I can think of right now. Good luck.

Thanks for your post. The machines giving the problems are all 8gb, so I am hoping that isn’t the problem. I’ll have a look at the application load settings, but to be honest, I am pretty sure it defaults to 20 minutes which is far far more time than it takes to load the scene. I suppose it could be network bottlenecks, but we are all gigabit networked here with cat6 cables, and so it should be fine.

The main thing is it seems to be these dual quad-core machines only and not the dual dual-core which have just 4gb ram. Very odd. I hate the way that BB is sooo difficult to troubleshoot.

Are you able to remote desktop into the nodes to watch what they do when they pick up a job? That might give you some more insight into the problem. Also check the vray log file on the computer. It may just be that BB is throwing a generic error for something vray is doing.

We’re using a pretty basic gigabit network here with cat5e wiring, so you shouldn’t be having problems there.

Oh, and a commonly recommended solution to a lot of backburner woes is to make sure you have the full max package installed on the manager machine. It’s stupid, I know, but it often fixes problems. Why? I don’t know. I had some problems a while back and I completely uninstalled backburner on the manager and re-installed the whole max package and everything was fine after that. Something to try.

I just thought of something else. Are these new machines? Have they successfully rendered anything yet? I know that max takes a while to load right after its first installed. This could be causing a time-out if the scene is large and max is having trouble launching anyway. Maybe try sending a simple job out to make sure everything is working right.

They are newish, and so they have rendered quite a bit of stuff. Will give the complete reinstall an idea.

Another thing I have been noticing is that when I remote desktop into the offending machines (which I can do so the network is not completelky buggered!), the mapped netwrk drives (a work drive, an admin drive, a textures drive etc) are marked as being disconnected. Double-clicking on them doesn’t re-establish the link which is odd. Now I can’t be sure that this happens every time we have one of these BB problems, but certainly over the last couple of days it seems pretty consistent.

This suggests it could be some network issue, but I don’t really know what. I can RDP into the machines, so the network hasn’t completely died, as I have said, so it just seems to be the link from the machines to the mapped drives on the Win2003 Server machine. This is odd. What could be going on there. I have, this morning, updated the network drivers on these machines - its an Intel PRO/1000 EB NIC (2 of them in each machine, but we only use the one), so I have yet to see if this helps.

We used to have the occasional mapped drives issue many years ago, but since moving to a proper Domain with a ‘serious’ server, the problems seemed to go away. Even back when we had the problems, double-clicking on the mapped drives seemed to re-establish the link.

Any suggestions?

Do you use Symantec Endpoint Security (anti virus) by any chance? There were some nasty bugs in the original release that would cause network drives to lose their mapping and need a reboot to fix it.

We’ve been on a domain since day one but we had one machine that would lose its mappings for some unknown reason. We eventually replaced the onboard NIC with a PCI NIC and the problem went away. So, it could be a driver issue. It could also be a NIC configuration issue. Bad cable? Remote desktop is pretty tolerant of a bad connection as it will often allow a connection to die and reconnect without you even knowing it. However, BB may not be so tolerant.

I’m not a network specialist by any means but heres a list of things to look into for possible problems/solutions based on problems we’ve run into here:

Disabling connection AutoSense (the thing that knows when a cable is plugged in or unplugged)

Manually specifying a connection speed instead of Auto Negotiate

Disabling SMB signing of packets (we found that disabling it greatly decreased latency and solved other problems we were having)

Make sure that if you system goes into ‘Offline Mode’, that it automatically goes back ‘Online’. I seem to remember these being in the domain or computer policy settings on the server.

Check firewall settings or temporarily disable the firewall on the offending nodes to see if the problem goes away.

Hope this helps.

We had the same issue a while ago - just scanned through so it may have been mentioned :
make sure that you have the same version of backburner installed on all machines. Max2009 does not install backburner if one is allready installed - so we ended up with machines with different versions of backbuner. We sorted this out and all our problems went away.